Model selection
Reproducible comparisons across providers and open models, on your use cases — not on a paid benchmark.
We evaluate models and architectures with client data —in Spanish, French and English— before committing cost or trust. What fails evaluation doesn't ship.
Slash AI Lab is Slash Digital Lab's evaluation pillar: we test models —OpenAI, Anthropic, Gemini, DeepSeek, GLM, Qwen, Mistral and local open source— on the client's real prompts before committing cost or private data.
Lab rule: no model is chosen by hype. It's chosen by evaluation on the client's prompts, cost, data residency and observed failures.
Evaluation axes: business Spanish, French, code, agents, cost, latency and privacy.
Reproducible comparisons across providers and open models, on your use cases — not on a paid benchmark.
Your operation's Spanish and French are evaluated as hard as English; nuances matter.
MCP, function calling, memory and human oversight: we measure what breaks when the agent errs.
Inference budget, p95 response time and data residency: cloud, VPC or local inference.
We gather real or representative prompts and success criteria with your team.
Candidates run in cloud or local; we measure quality, cost and latency.
Results per axis with examples: a clear, reproducible recommendation.
The market moves every quarter; when it does, we re-run.
| Evaluation | 30 fixed prompts, 5 engines, 3 languages (methodology published in /insights/). |
|---|---|
| Data residency | Standard cloud, VPC or local inference depending on compliance. |
| Report | Comparison per axis: language, code, agents, cost, latency, failures. |
| Reproducibility | Prompts, versions and criteria documented to repeat the run. |
| Cadence | Quarterly re-evaluation or on relevant market moves. |
| Publication | Findings open in /insights/ and /research/; client prompts never. |
The ones that can earn a slot: OpenAI, Anthropic, Gemini, DeepSeek, GLM, Qwen, Mistral and local open source. The list moves with the market.
Yes: method and findings go to /insights/ and /research/. Client prompts are never published.
Recommended, not mandatory. For sensitive-data or minority-language projects, evaluation avoids months of drift.
Yes: open models in your VPC or own hardware when compliance requires it.
Yes: one focused run with your prompts and a reproducible report.
Applied AI for real operations: agents, RAG, vision, voice and automation.
Learn more→ Fine-tuningWhen to tune the model and when to give it knowledge. Judgement, not hype.
Learn more→ AgentsAgents with tools, memory and human oversight: support, ops and back office.
Learn more→ ModelosEvaluation in Spanish and French. Production criteria.
Learn more→Tell us your challenge. We reply within 24 hours with an honest first read: if we can help, we'll say how; if not, we'll say who can.
I reply personally. No endless forms, no canned replies.