Slash AI Lab · September 2026

How we choose a model

A model isn't chosen by hype. It's chosen by evaluation on client prompts, cost, data residency and observed failures. Here's how we run it, openly.

30 fixed prompts5 engines3 languages
In short

We evaluate the candidates —OpenAI, Anthropic, Gemini, DeepSeek, GLM, Qwen, Mistral and local open source— against the client's real prompts in Spanish, French and English. A model isn't chosen by fashion; it's chosen by observed failures and defensible cost.

The frame

The frame before the model

Five constraints decide; the ranking follows.

Choosing a model in 2026 doesn't start with a leaderboard: it starts with your constraints. A model that tops an English benchmark can fail your operation's Spanish, cost twice per token, or live where your compliance forbids it.

The Slash AI Lab frame: language, code, agents, cost, privacy. Five axes, measured with the client's real prompts, against the best generic model available as baseline.

Method

How the evaluation runs

Reproducible by design: anyone can re-run it.

We gather 30 fixed prompts per client — real business tasks, not toys — and run each candidate in cloud and, where applicable, in local inference.

Each run measures four things: quality (against agreed criteria), observed failures (what breaks when things go wrong), inference cost and p95 latency. The output is a per-axis comparison with a reproducible recommendation.

  • Prompts: 30 fixed per client, versioned.
  • Engines: 5 in the current run (method published in the lab).
  • Languages: business Spanish, French, English.
  • Output: per-axis comparison with a reproducible recommendation.
Failures

Failures outweigh successes

Deployments are decided by what breaks, not by hits.

Success benchmarks reward the easy. Deployments are decided by what happens when the model errs: does it answer with false confidence, invent sources, ignore the client's language?

The evaluation therefore includes hostile cases: adversarial prompts, dirty data, broken tools. If the model fails badly, it goes back to design — regardless of its ranking.

Limitations

What this study is NOT

Declared limitations, no smoke.

It's not a paid or sponsored ranking; it doesn't cover every vendor; and it doesn't replace an evaluation on your own prompts — it's the public base of the method Slash runs on every project.

Key takeaways

  • The frame: language, code, agents, cost, privacy — discussed in that order.
  • 30 fixed prompts per client, versioned and reproducible.
  • Observed failures outweigh benchmark hits.
  • Local inference when compliance requires it.
  • Re-evaluated quarterly or on market moves.
Frequently asked questions

The questions we hear often

Which engines ran?

OpenAI, Anthropic, Gemini, DeepSeek, GLM, Qwen, Mistral and local open source; the exact list is documented in the lab.

Why 30 prompts?

It's the minimum that separates luck from consistency on real tasks; more rarely changes the conclusion.

Why Spanish and French?

Because that's where models break: business Spanish and French are Slash working languages, not extras.

Is it a paid ranking?

No. The methodology, prompts and limitations are public.

How often is it re-run?

Quarterly, or on any market move that matters.

Talk to Slash

Let's put it in production

Tell us your challenge. We reply within 24 hours with an honest first read: if we can help, we'll say how; if not, we'll say who can.

I reply personally. No endless forms, no canned replies.