RAG in production
Indexing, re-ranking, citations and answer evaluation over your real documents.
RAG when knowledge must stay current and citations matter. Fine-tuning when style, language or a narrow skill doesn't fit in context. Hybrid when both are true.
There's no magic recipe — there are constraints. If your data changes daily, RAG. If tone or format must be yours, fine-tuning. If compliance requires keeping knowledge in-house, local inference with open models.
Read the full criterion in Fine-tuning vs RAG and how we choose models in How we choose a model.
The cheapest path that meets the quality bar; the rest is consultant margin.
Indexing, re-ranking, citations and answer evaluation over your real documents.
LoRA adapters for style, format or narrow skills; no money burned on what RAG already solves.
Test sets from your real cases and cost/quality comparison across routes.
Data residency, local inference or VPC when compliance requires it.
Data, languages, budget and compliance: the frame before the model.
We test the best generic model; that's the yardstick everything else must beat.
RAG and/or adapters over the baseline with comparable metrics and cost per run.
The cheapest route that meets the bar, documented. If none exists, we say so.
| Constraint analysis | Data, language, budget and compliance, in writing. |
|---|---|
| Baseline | The best generic model evaluated on your cases. |
| Experiment | RAG and adapters with comparable metrics and cost. |
| Recommendation | The chosen route with estimated cost and risks. |
| Implementation | A production pipeline if you decide to continue. |
| Documentation | The criterion, documented and shareable. |
When the problem is missing or changing knowledge — RAG solves that cheaper and safer.
Yes, including local inference with GLM, Qwen, Mistral and others when compliance requires it.
It's scoped in sprints; discovery sets the size. Write to us with your case.
We explicitly evaluate your business Spanish; it's one of the lab's axes.
When style must be yours and knowledge must be current: adapters for tone, RAG for the data.
Our lab where we evaluate models and publish the method.
Learn more→ AIApplied AI for real operations: agents, RAG, vision, voice and automation.
Learn more→ RAGWhen RAG, when fine-tune, when hybrid.
Learn more→ ConsultingDiagnosis and architecture before building. Actionable roadmaps.
Learn more→Tell us your challenge. We reply within 24 hours with an honest first read: if we can help, we'll say how; if not, we'll say who can.
I reply personally. No endless forms, no canned replies.