Everything MARK-17 does.

Organised by area, in plain language. If you want the short version of why any of this matters, start with the MARK-17 overview.

998
Scoring dimensions
Ways a model is measured, not just whether it answered
254
Tests
Grouped into 11 domains, from reasoning to adversarial safety
13
Providers
Local runtimes and cloud APIs, including custom endpoints

What MARK-17 can benchmark.

Models running on your own hardware, models from cloud providers, or both compared against the same tests.

On your hardware

  • Built-in llama.cpp server
  • Ollama
  • LM Studio
  • LocalAI

Cloud providers (your own API keys)

  • OpenAI
  • Anthropic
  • Google Gemini
  • xAI (Grok)
  • Groq
  • DeepSeek
  • Mistral
  • Together AI
  • Any OpenAI-compatible endpoint

By area.

Benchmarking engine

  • Evaluation engine with composite score and per-domain breakdown
  • Per-test scoring, pass/fail results, and anomaly detection
  • Empirical context-window validation
  • Fixed-seed runs recorded with a hardware fingerprint
  • Latency and throughput metrics with hardware-efficiency analysis
  • Re-run a benchmark or retry only the tests that failed or timed out
  • Scheduled benchmark runs that execute without you sitting there

Choosing models

  • Guided wizard with live monitoring and in-wizard cost estimates
  • Recommendation engine considering hardware fit, intent, and budget
  • Local models via llama.cpp, Ollama, LM Studio, CUDA, and Vulkan
  • Cloud models using your own provider API keys
  • Full model catalog with discovery, filtering, and VRAM fit analysis
  • Model queue for lining up work across several models

Test coverage

  • Reasoning, Coding, and Chat
  • Deployment Risk, Adversarial Safety, and Tool Calling
  • Multimodal, Multi-Turn Adversarial, and Agentic
  • Agentic Email and Context Retention
  • Per-run enabling and disabling of individual tests
  • Custom test authoring, and custom suites saved and version-locked

Comparison and history

  • Side-by-side comparison, domain by domain
  • Mix local and cloud models in the same comparison
  • Saved session history, stored locally on your machine
  • Benchmark run archive you can go back through at any point

Reports and exports

  • Full professional report with cover, charts, latency, cost, and appendices
  • Comparison reports, executive briefs, and short report briefs
  • CSV export for spreadsheet analysis
  • JSON full-audit export
  • Standalone cost-report PDF
  • MBX export — a tamper-evident verification artifact
  • Batch export for producing many reports at once
  • Report branding, including custom logo and a "Prepared For" client name

Cost and operating projections

  • Pre-run cost estimate before you spend anything
  • Cloud API token cost estimates per provider
  • Local electricity cost estimates from your own power assumptions
  • Daily, monthly, and annual operating projections
  • Configurable electricity rate and system power settings
  • Cost warnings before an expensive run starts

Consulting workflow

  • Advisor: reads your benchmark evidence and guides discovery review through to implementation planning
  • Self-Proving Task Agent: prove one specific client task actually works before you promise it
  • Clients and projects, with saved intakes and an Advisor archive
  • Cost estimator producing planning-level effort and return ranges
  • Client proposals and staged implementation plans
  • Decision packages for organisations that need a formal write-up

Hardware safety and security

  • Thermal protection, cooldown, and VRAM monitoring
  • Tamper detection and encrypted local key storage
  • Local-first data handling
  • Component manager, on-device judge, and first-run setup
  • Built-in plugin management with integrity validation

Audit and compliance

  • Complete audit trail with timestamps and actors
  • Tamper-evident integrity using a SHA-256 hash chain
  • Chain-of-custody reporting
  • Compliance audit export
  • SIEM-compatible exports in JSON and CSV
  • Integrity verification of the audit log

A note on licenses: This is what MARK-17 can do, not a list of what every license includes. Some advanced capabilities depend on your license. If a specific one matters to your decision, ask us before you buy and you will get a straight answer.

What you get during the trial: the 14-day trial is a real evaluation — Advisor and the full consulting workflow, the benchmarking engine, and every test suite, with PDF reports, printing, and MBX evidence packages. Trial reports are watermarked, and the formats that would remove that watermark — clean or white-label output and CSV, JSON, and audit exports — are what a license unlocks. MBX is the exception: it is provenance and tamper-evidence material, not a finished deliverable. External developer interfaces are not part of it. When the trial ends, productive work stops — no new runs, Advisor work, reports, or exports — but nothing is deleted: everything you produced stays on your machine and stays readable. See the full breakdown.

The paid classes are Business and Enterprise. Which capabilities sit on each side of that line is not published yet — if it affects a purchasing decision, ask us directly.

Where you actually work.

MARK-17 is a working application, not a single screen with a Run button. These are the areas you move between.

Advisor

The consulting workflow, from the first client conversation to an implementation plan.

Clients and Projects

Who the work is for, and what each engagement covers.

Test Suites

The built-in suites, plus any custom suite you build yourself.

Model Queue and Scheduled Runs

Line up benchmark work and let it run when you are not watching.

Reports

Benchmark reports, proposals, implementation plans, and evidence packages in one place.

Models

The catalog, your local models, and the cloud models you have connected.

Saved Intakes and Advisor Archive

Earlier discovery sessions and finished Advisor work, kept for reference.

Audit Log

What happened, when, and who did it — with tamper-evident integrity.

The actual application.

Select any image to view it full size.

MARK-17 dashboard with completed benchmark sessions
The MARK-17 wizard in simple mode, describing a task in plain English
The MARK-17 wizard in advanced mode with full step-by-step control
Model selection in MARK-17 showing hardware fit for each model
A benchmark running with live hardware telemetry
A completed benchmark session with summary scores
Local models detected from Ollama and LM Studio
The MARK-17 model catalog with filtering
A generated MARK-17 report with inference metrics

Don't pay for another AI model until you know it can do the job.

See how MARK-17 works end to end, or read the documentation before you decide anything.