Find the right AI for the job — local or cloud.
Tell MARK-17 what you need AI to do. It helps you choose models, test them, compare the results, and see which options actually work best for your needs.
Then it keeps going: the evidence becomes a report you can defend, a plan for putting it into the business, and an estimate of what it costs.
It is not just a benchmark tool. It is a proof-and-evidence system: it helps you choose models, test them, compare the outputs, understand why one failed, and produce evidence that stands up when somebody asks how you decided. Built for consultants, builders, technical teams, and businesses that have to be right about this.
14-day trial, no payment details required. Or see pricing and licensing · read the docs.
It does not stop at a score.
Plenty of tools will hand you a benchmark number. MARK-17 starts from a real business task and carries it all the way to a costed plan — the evaluation, the evidence, the decision, and the documents you need to get it approved and built.
- What you start with A real business task
- “We want AI to draft replies to customer email — without inventing facts.”
- Written in plain English. No benchmark expertise required.
- What MARK-17 does Evaluation, evidence, decision
- Evaluation — the same tests against every candidate
- Evidence — scores, speed, and hardware behaviour, kept
- Decision — the option that fits the job
- What you walk away with Documents you can act on
- Proof report — evidence anyone can check before they approve it
- Implementation plan — how it would be put into the business
- Cost estimate — what the approach is likely to cost to run
Diagram of the MARK-17 workflow — not a screenshot of the application.
Most AI decisions are made on the wrong information.
Companies pick AI models based on popularity, marketing, a leaderboard score, an assumption, or how the model performed on somebody else's workload. None of those tell you how a model handles your work.
The cost shows up later: paying for capability you never use, discovering the cheap option cannot do the job, or committing to hardware that turns out to be wrong for the models you actually need.
MARK-17 tests models against the work that actually matters to you — and keeps the results so you can show your reasoning to anyone who needs to approve it.
Anyone who has to answer “which AI should do this?” and be right.
The question is the same whether you are advising a client, building a product, or deciding what your own company should run on. What differs is who you have to convince afterwards.
Consultants and agencies
You are advising someone else, and your recommendation has to hold up when they ask why. MARK-17 gets a dedicated workflow for this further down the page.
People becoming AI consultants
You know the work is there. What you do not have yet is a way to show a stranger that your recommendation is more than an opinion. That is the gap this closes.
Builders and inventors
You are putting AI inside something you are making, and the wrong model shows up later as a product that does not work well enough.
Technical teams
You have to pick something, deploy it, and live with the consequences. Popularity is not evidence and a leaderboard is not your workload.
Businesses choosing for themselves
You are spending real money on AI and want to know whether the expensive option is actually better at the job you need done.
People weighing local against cloud
The answer depends on your hardware, your data, and your budget. It is a testable question, so test it.
Anyone who has to justify the choice
If someone above you has to approve it, an opinion is a weaker position than a report with the runs behind it.
That list is not meant to be exhaustive. If you have a real job and more than one AI that might do it, the question MARK-17 answers is yours too.
From a business task to a costed plan.
Seven stages. You stay in control at every one of them, and you can stop at any of them — plenty of people only need the first four.
- 1
A real business task
Describe the work you need AI to do, in plain English.
- 2
Evaluation
Candidate models are tested against that task under the same conditions.
- 3
Evidence
Scores, speed, and hardware behaviour are recorded and kept.
- 4
Decision
You pick the option the evidence actually supports.
- 5
Proof report
A document that shows your reasoning to whoever has to approve it.
- 6
Implementation plan
How the chosen approach would be put into the business.
- 7
Estimate
What that approach is likely to cost to build and to run.
Test local AI, cloud AI, or both.
Sometimes AI running on your own hardware is the best fit. Sometimes a cloud model is the better answer. Sometimes a business needs both. MARK-17 helps you test the available options instead of guessing which way to go.
AI on your own hardware
MARK-17 ships with a managed built-in llama.cpp runtime, so you can benchmark local models without installing anything else first. Already using Ollama, LM Studio, or LocalAI? Connect them instead — MARK-17 detects the models you have and reports how they fit the machine you are testing on.
Useful when data residency, running costs, or independence from a provider matter.
AI from cloud providers
Connect supported cloud providers using your own API keys and benchmark those models against exactly the same tests.
Useful when you need capability that is impractical to run yourself, or want to know whether paying per call beats buying hardware.
Repeatable tests, not one-off impressions.
Every model is tested under the same conditions, so the comparison means something. MARK-17 watches what happens while the tests run — memory use, hardware load, timing — and records it alongside the scores.
The result is a like-for-like comparison you can repeat later to check whether a model has changed.
- TTFT
- How quickly the model starts responding.
- TPOT
- How quickly it produces each piece of the answer.
- TPS
- Overall throughput — how much work it gets through.
- E2E
- Total time from question to finished answer.
998 scoring dimensions across 254 tests.
Tests are grouped into 11 domains, each covering a different kind of capability. Every test carries multiple scoring criteria — measuring not just whether a model answered, but how well.
Reasoning
30Logic, math, probability, constraint satisfaction, meta-reasoning
Coding
35Algorithms, data structures, concurrency, systems programming, edge cases
Chat
25Conversation quality, format compliance, creativity, empathy, bias detection
Multimodal
30Image understanding, OCR, charts, diagrams, medical imagery, satellite
Deployment Risk
28Safety refusals, prompt injection defense, PII handling, jailbreak resistance
Adversarial Safety
30Role-play bypass, authority injection, sycophancy, obfuscated attacks, instruction conflicts
Tool Calling
33Function calling accuracy, parallel/sequential, error recovery, fault tolerance
Agentic
27Goal decomposition, multi-agent coordination, state management, autonomous troubleshooting
Multi-Turn Adversarial
8Gradual escalation, persona persistence, language switching across turns
Agentic Email
1Real-world email inbox management task
Context Retention
7Needle-in-haystack from 8K to 1M tokens
Results somebody else can check.
MARK-17 can package a whole benchmark session into a single MBX file sealed with a SHA-256 content hash. If anything in that file is changed afterwards, the hash stops matching and the change is detectable.
The verifier is open source and published separately, so the person receiving your results does not have to take your word for it — or ours.
What this does and does not mean
MBX gives you tamper evidence: proof that a results file has not been edited since it was exported. That is genuinely useful when results travel between machines and people.
It is not a legal certification, a government approval, or forensic proof that a benchmark was run honestly in the first place. We say so plainly because overstating it would defeat the purpose.
Understand what is needed, test the options, show the evidence, and turn the findings into a plan.
Benchmarking is only part of the work. Advisor is the ten-step consulting workflow that connects the rest of it — from the first conversation about the problem to a staged plan for doing something about it.
It is at its strongest when you are advising someone else and your recommendation has to survive scrutiny. It works just as well when the person you have to convince is your own management.
- 01 Advisor Capture the problem, the outcome wanted, the constraints, and where it has to run.
- 02 Discovery Review Check there is genuinely enough to work with before going further.
- 03 Advisor Analysis See what was understood from the intake, and what is still missing.
- 04 Model Plan Decide which models are worth evaluating, and which to rule out.
- 05 Benchmark & Evidence Run the tests, then qualify what the results do and do not prove.
- 06 Solution Direction Form a recommendation from the evidence — or say plainly that the evidence is not there yet.
- 07 Estimator Planning-level ranges for cost, effort, and likely return.
- 08 Proposal A client-facing proposal drafted from the work you just did.
- 09 Implementation Plan Staged delivery, risks, who owns what, and where a person has to review.
- 10 Reports Everything stored exactly as it was saved, ready to reopen.
Advisor reads and uses the benchmark evidence you have gathered. It does not start benchmark runs on its own — you decide what gets tested and when. Step six will refuse to hand you a confident recommendation the evidence does not support, which is the point of it.
Prove one task actually works before you promise it.
Most engagements come down to a single job somebody wants AI to take over. The Task Agent is where you pin that job down and find out whether it genuinely holds up — before it becomes a commitment in a proposal.
- ✓ Define one task, precisely, and write down what "working" means.
- ✓ Check whether it is actually a sensible thing to hand to AI.
- ✓ Record real cases that represent the work honestly.
- ✓ Run hand-operated proof runs and keep what happened.
- ✓ Produce a proof summary from those runs.
- ✓ Feed the proof straight into your estimate.
The result is a provider-neutral starter kit for that engagement — evidence that the task works, not a promise that it will.
The work leaves the application.
Reports generated during the trial are watermarked — MBX excepted, as provenance and tamper-evidence material. A paid license is what unlocks the clean, white-label, and raw-data formats — and, once the trial expires, what lets you produce anything new at all.
Comparison reports
Models side by side on the same tests.
PDF reports
Presentation-ready documents with summaries.
Executive brief
The short version, written for whoever signs off.
Technical evidence report
The long version, for the people who will ask how you know.
Implementation blueprint
What actually gets built, in what order, with the risks named.
Client proposal
A proposal drafted from the evidence rather than from scratch.
Decision package
A formal write-up for organisations whose process requires one.
CSV and JSON
Raw data for your own analysis or tooling.
MBX packages
Tamper-evident evidence files others can verify.
Cost report
What the approach costs to run — daily, monthly, and annually.
The model that wins on quality is not always the one you can afford.
A score tells you which model is better. It does not tell you what running it every day for a year does to your budget. MARK-17 works that out alongside the benchmark.
Before you spend anything
The wizard estimates what a run will cost before it starts, and warns you when a run is about to be expensive.
Cloud costs, per provider
Token cost estimates for the cloud models you are testing, using the providers you actually connected.
Local costs, from your own power
Electricity cost estimates based on your rate and your system, because running a model on your own hardware is not free either.
Daily, monthly, and annual
Operating projections over real time horizons, so the comparison is between running costs rather than between benchmark numbers.
It does not need you sitting there
Benchmarks take time. You can queue work across several models and schedule runs to happen without you watching, then come back to finished results, a full history you can re-open, and an audit log of what happened and when.
A note on licenses: This is what MARK-17 can do, not a list of what every license includes. Some advanced capabilities depend on your license. If a specific one matters to your decision, ask us before you buy and you will get a straight answer. See trial and licensing.
Local-first, explained honestly.
MARK-17 runs on your machine and keeps your prompts, model outputs, and benchmark results in local storage. We do not collect your benchmark content.
Local-first does not mean permanently offline
Some things genuinely need the internet, and we would rather be straight about which:
- • License activation and updates.
- • Downloading models you choose to install.
- • Any cloud AI provider you connect — your test prompts go to that provider, under their terms.
- • Optional diagnostics, which can be turned off.
Full detail is in the Privacy Policy.
What you need to run it.
Minimum
- OS Windows 10 or 11 (64-bit)
- CPU 64-bit processor
- RAM 8 GB
- Disk 2 GB free space
Recommended for local models
- OS Windows 11 (64-bit)
- CPU Modern multi-core processor
- RAM 16 GB or more
- GPU NVIDIA with 8 GB or more VRAM
- Disk SSD storage
On platforms: Windows 10 and 11 (64-bit) is the proven customer platform today, and the only one with a released build. macOS and Linux are not available. If you need either, tell us — it counts toward what gets built next, but we are not going to put a date on it before we can stand behind one.
No GPU is required to run MARK-17 or to benchmark cloud models — that work happens on the provider's machines. Local models will run CPU-only if you have enough system RAM; a GPU simply makes them much faster and lets you run larger ones.
Fourteen days with the real product. Your work is yours either way.
The trial is a real evaluation, not a crippled demo — you get the working application, including Advisor. The line is drawn at output: the reports it produces are watermarked, and the formats that would remove that watermark are what a license unlocks — MBX aside, which is evidence material rather than a finished deliverable. When it ends, productive work stops — but nothing is deleted and nothing is taken away from you.
Trial
- • Starts at first launch and activation.
- • No payment information required.
- • Advisor and the full consulting workflow.
- • The benchmarking engine and every test suite.
- • PDF reports and printing — watermarked.
- • MBX evidence packages included.
- • The reports it produces carry a watermark.
- • No clean, white-label, or data exports.
- • External developer interfaces not included.
Trial Expired — Read-Only
- • Productive work stops. Nothing uninstalls.
- • No new benchmark runs or scheduled runs.
- • No new Advisor or Task Agent work.
- • No new reports and no new exports.
- • Your projects, results, and evidence stay put.
- • You can still open and read all of it.
- • A license turns productive work back on.
Business or Enterprise
- • Two paid license classes.
- • Business permits paid client work.
- • Consultants and agencies are covered by Business.
- • Perpetual — the version you bought keeps working.
- • 12 months of update coverage included.
- • Lapsed coverage does not disable your copy.
On price: pricing is being revised ahead of the next release, so we are not publishing a figure that changes next month. Ask and you will be told where it stands. Full license terms are in the Terms of Service and EULA.
Questions people actually ask.
What does MARK-17 actually do?
It helps you pick an AI model for a specific job and prove the choice was right. You describe the work, MARK-17 helps you find models worth testing, runs the same tests against each one, and shows you how they compared on quality, speed, and hardware behaviour.
Is MARK-17 a chatbot?
No. It is a desktop application for testing and comparing AI models. It does not answer your questions — it measures how well other AI models answer them.
Can it test AI models running on my own hardware?
Yes. MARK-17 ships with a managed built-in llama.cpp runtime, so you can benchmark local models without setting up anything else first. If you already use Ollama, LM Studio, or LocalAI, you can connect those instead and MARK-17 detects the models you have installed.
Can it test cloud AI models?
Yes. You can connect supported cloud providers with your own API keys and benchmark those models the same way.
Can I compare a local model against a cloud model?
Yes — that is one of the main reasons to use it. Sometimes local is the better fit, sometimes cloud is, and sometimes a business needs both. MARK-17 lets you test rather than guess.
Does my data leave my computer?
MARK-17 is local-first: your prompts, model outputs, and benchmark results are stored on your machine, not on our servers. Local-first does not mean permanently offline — activation, updates, downloading models, and any cloud AI provider you choose to connect all use the internet. If you benchmark a cloud model, your test prompts necessarily go to that provider.
Do I need a GPU?
No. You do not need a GPU to run MARK-17, and you do not need one to benchmark cloud models. Local models can run CPU-only provided you have enough system RAM — they are just slower. A GPU with enough VRAM makes local benchmarking substantially faster and lets you run larger local models. MARK-17 reports how each model fits your hardware, so you can see what your machine will actually manage.
Which operating systems are supported?
Windows 10 and 11 (64-bit) is the proven customer platform today, and the only one with a released build. macOS and Linux are not available and we are not going to imply otherwise. If you need either, say so and it counts toward what gets built next.
What is MBX?
MBX is the export format MARK-17 uses to package a benchmark session into a single file with a SHA-256 content hash. If the file is altered afterwards, the hash no longer matches. An open-source verifier lets anyone check a file without installing MARK-17.
What is Advisor?
Advisor is the guided consulting workflow inside MARK-17. It reads and uses the benchmark evidence you have already gathered to help you interpret the results and decide what to do about them, connecting discovery, analysis, a model plan, cost estimates, proposals, and implementation planning. Advisor does not launch benchmark runs itself — you decide what gets tested and when. Advisor is available in full during the 14-day trial, and stops when the trial expires unless you have licensed MARK-17.
Can I export results?
Yes. Comparison reports, PDF reports, executive briefs, technical evidence reports, implementation blueprints, client proposals, CSV and JSON data, cost reports, and MBX evidence packages. Reports produced during the trial are watermarked — MBX excepted, as provenance and tamper-evidence material — and the clean and raw-data formats that would remove that watermark are unlocked by a license. Once the trial expires, no new exports are produced at all until you license MARK-17 — though everything you already exported or saved stays yours. Some of the more advanced export and reporting formats depend on your license — ask us if a specific one matters to your decision.
Is there a trial, and what happens when it ends?
Yes — 14 days, starting when you first launch and activate MARK-17, with no payment information required. It is meant to be a real evaluation: you get Advisor and the full consulting workflow, the benchmarking engine, and every test suite, and you can generate PDF reports, print them, and produce MBX evidence packages. The line is drawn at output: the reports the trial produces are watermarked, and the formats that would remove that watermark — clean PDF, white-label branding, and CSV, JSON, and audit exports — are what a license unlocks. MBX is the deliberate exception, because it is provenance and tamper-evidence material rather than a finished deliverable. External developer interfaces are not part of the trial. When the 14 days are up, productive work stops: no new benchmark runs, scheduled runs, Advisor or Task Agent work, reports, or exports. Nothing is uninstalled and nothing is deleted — the projects, results, and evidence you produced stay on your machine and stay readable in the application, so you can go back over what you found. Activating a license turns productive functionality back on.
Which license do I need if I work for clients?
The two paid classes are Business and Enterprise, and both are perpetual with 12 months of update coverage included. A Business license explicitly permits paid client work, including consultants and agencies delivering benchmark evidence and reports to the businesses they serve — you do not need Enterprise merely because you have clients.
What does it cost?
Pricing is being revised ahead of the next release, so rather than publish figures that are about to change we would rather you contact us directly — email sales@aibenchlab.com and we will tell you exactly where things stand.
Consulting on AI, and need more clients?
MARK-17 proves the AI works. LYDIA-12 helps you find the businesses that need it — and arrive already knowing what is wrong with theirs. Between them they cover the path from wanting this work to being paid for it.
Neither product requires the other. LYDIA-12 is still under development; MARK-17 is available now.
See what LYDIA-12 does →Stop guessing which AI to use.
Read the documentation, see what gets tested, or ask us a question.