How we choose a model
A model isn't chosen by hype. It's chosen by evaluation on client prompts, cost, data residency and observed failures. Here's how we run it, openly.
We evaluate the candidates —OpenAI, Anthropic, Gemini, DeepSeek, GLM, Qwen, Mistral and local open source— against the client's real prompts in Spanish, French and English. A model isn't chosen by fashion; it's chosen by observed failures and defensible cost.
The frame before the model
Five constraints decide; the ranking follows.
Choosing a model in 2026 doesn't start with a leaderboard: it starts with your constraints. A model that tops an English benchmark can fail your operation's Spanish, cost twice per token, or live where your compliance forbids it.
The Slash AI Lab frame: language, code, agents, cost, privacy. Five axes, measured with the client's real prompts, against the best generic model available as baseline.
How the evaluation runs
Reproducible by design: anyone can re-run it.
We gather 30 fixed prompts per client — real business tasks, not toys — and run each candidate in cloud and, where applicable, in local inference.
Each run measures four things: quality (against agreed criteria), observed failures (what breaks when things go wrong), inference cost and p95 latency. The output is a per-axis comparison with a reproducible recommendation.
- Prompts: 30 fixed per client, versioned.
- Engines: 5 in the current run (method published in the lab).
- Languages: business Spanish, French, English.
- Output: per-axis comparison with a reproducible recommendation.
Failures outweigh successes
Deployments are decided by what breaks, not by hits.
Success benchmarks reward the easy. Deployments are decided by what happens when the model errs: does it answer with false confidence, invent sources, ignore the client's language?
The evaluation therefore includes hostile cases: adversarial prompts, dirty data, broken tools. If the model fails badly, it goes back to design — regardless of its ranking.
What this study is NOT
Declared limitations, no smoke.
It's not a paid or sponsored ranking; it doesn't cover every vendor; and it doesn't replace an evaluation on your own prompts — it's the public base of the method Slash runs on every project.
Key takeaways
- The frame: language, code, agents, cost, privacy — discussed in that order.
- 30 fixed prompts per client, versioned and reproducible.
- Observed failures outweigh benchmark hits.
- Local inference when compliance requires it.
- Re-evaluated quarterly or on market moves.
The questions we hear often
Which engines ran?
OpenAI, Anthropic, Gemini, DeepSeek, GLM, Qwen, Mistral and local open source; the exact list is documented in the lab.
Why 30 prompts?
It's the minimum that separates luck from consistency on real tasks; more rarely changes the conclusion.
Why Spanish and French?
Because that's where models break: business Spanish and French are Slash working languages, not extras.
Is it a paid ranking?
No. The methodology, prompts and limitations are public.
How often is it re-run?
Quarterly, or on any market move that matters.
Where to next
Slash AI Lab
Our lab where we evaluate models and publish the method.
Learn more→ Fine-tuningFine-tuning & RAG
When to tune the model and when to give it knowledge. Judgement, not hype.
Learn more→ AIArtificial intelligence
Applied AI for real operations: agents, RAG, vision, voice and automation.
Learn more→ ResearchResearch
We publish our method: models, RAG, pentest and LLM visibility.
Learn more→Let's put it in production
Tell us your challenge. We reply within 24 hours with an honest first read: if we can help, we'll say how; if not, we'll say who can.
I reply personally. No endless forms, no canned replies.