Fine-tuning & RAG

Judgement, not recipes

RAG when knowledge must stay current and citations matter. Fine-tuning when style, language or a narrow skill doesn't fit in context. Hybrid when both are true.

Default RAGWhen it fits Adapters · LoRARule Evaluation > hype
In short

There's no magic recipe — there are constraints. If your data changes daily, RAG. If tone or format must be yours, fine-tuning. If compliance requires keeping knowledge in-house, local inference with open models.

Read the full criterion in Fine-tuning vs RAG and how we choose models in How we choose a model.

What we deliver

What we build

The cheapest path that meets the quality bar; the rest is consultant margin.

01

RAG in production

Indexing, re-ranking, citations and answer evaluation over your real documents.

02

Purposeful fine-tuning

LoRA adapters for style, format or narrow skills; no money burned on what RAG already solves.

03

Honest evaluation

Test sets from your real cases and cost/quality comparison across routes.

04

Privacy

Data residency, local inference or VPC when compliance requires it.

How we work

How we decide and build

  1. Constraints

    Data, languages, budget and compliance: the frame before the model.

  2. Baseline

    We test the best generic model; that's the yardstick everything else must beat.

  3. Experiment

    RAG and/or adapters over the baseline with comparable metrics and cost per run.

  4. Decision

    The cheapest route that meets the bar, documented. If none exists, we say so.

Scope and deliverables

What the project delivers

Constraint analysisData, language, budget and compliance, in writing.
BaselineThe best generic model evaluated on your cases.
ExperimentRAG and adapters with comparable metrics and cost.
RecommendationThe chosen route with estimated cost and risks.
ImplementationA production pipeline if you decide to continue.
DocumentationThe criterion, documented and shareable.
Frequently asked questions

The questions we hear often

When is fine-tuning NOT the answer?

When the problem is missing or changing knowledge — RAG solves that cheaper and safer.

Do you work with open models?

Yes, including local inference with GLM, Qwen, Mistral and others when compliance requires it.

What does an experiment cost?

It's scoped in sprints; discovery sets the size. Write to us with your case.

Will results improve in Spanish?

We explicitly evaluate your business Spanish; it's one of the lab's axes.

When do you recommend hybrid?

When style must be yours and knowledge must be current: adapters for tone, RAG for the data.

Talk to Slash

Let's put it in production

Tell us your challenge. We reply within 24 hours with an honest first read: if we can help, we'll say how; if not, we'll say who can.

I reply personally. No endless forms, no canned replies.