AI agents & LLM integration

Language models put inside a product that has a job to do — not a chatbot bolted onto a homepage.

The gap between an impressive demo and a feature people use every day is almost entirely engineering. A model call is trivial. Knowing what to retrieve, how to bound the task, what to do when the model returns nonsense, how to keep costs predictable and how to keep your documents out of somebody else's training data — that is the work.

I build LLM features into real products: assistants that answer from your own material rather than from the open internet, retrieval pipelines over your own documents including entirely self-hosted models where the data cannot leave your infrastructure, and agent pipelines that carry a multi-step task through to a result instead of just answering a question and stopping.

I have built this into a product rather than into a demo — AI music generation inside a desktop application, where the interesting parts were not the model but everything around it: a generation is a job with states, not a function call; the completion callback lands on a server rather than on an app holding a connection open; and the result has to arrive somewhere a person was already working. That is the shape of most of this work.

What you get

Retrieval over your own content

Your documents, contracts, catalogue or knowledge base made answerable — with citations back to the source, so an answer can be checked rather than trusted.

Self-hosted where it matters

When the data genuinely cannot leave your infrastructure, open models running on your own hardware. Slower to build, and sometimes the only acceptable answer.

Agents that finish a task

Multi-step pipelines with tools, checkpoints and defined failure behaviour, so a task ends in a result or a clear failure — never in a confident hallucination presented as done.

Costs and limits you can predict

Token budgets, caching, model selection per step, and a fallback path when a provider is down. AI features have an ongoing bill, and it should not be a surprise.

How it works

  1. 01

    Find the task worth automating

    Usually not the one people first suggest. The best candidates are high-volume, tolerant of a review step, and currently done by a person reading and re-typing.

  2. 02

    Build the evaluation before the feature

    A set of real cases with known-good answers. Without it, "it seems better now" is the only available measure, and it is worthless.

  3. 03

    Ship it behind a human

    The first version proposes and a person approves. Confidence to remove that step should be earned from measured results, not assumed.

  4. 04

    Tighten cost and reliability

    Once it works, make it cheap and boring: caching, smaller models where they suffice, retries and graceful degradation.

Typically built with

GoTypeScriptVector searchSelf-hosted open modelsRAG pipelinesAgent tooling

Where I have done this

Common questions

Can we do this without sending our data to OpenAI or Anthropic?

Yes. Open models self-hosted on your own infrastructure are a genuine option, and for contracts, personnel files or anything under strict data protection they are often the only option that survives legal review. The trade-off is real: you get control and no per-token bill, in exchange for hardware, slower iteration and somewhat weaker models. I will tell you which side of that trade your use case falls on.

How do you stop it inventing answers?

Retrieval with citations, so every answer points at a source a person can open; bounded tasks rather than open-ended ones; and an evaluation set that catches regressions before your users do. You cannot reduce hallucination to zero, which is why the honest design keeps a human in the loop wherever a wrong answer would be expensive.

Is this worth it for a small company?

Sometimes not, and I will say so. If the task happens ten times a month, a person doing it is cheaper than a system that needs maintaining. The economics work when volume is high and the work is genuinely repetitive — document processing, classification, extraction, first-draft generation.

What does it cost to run?

Hosted models bill per token, so cost scales with usage and needs designing for — caching, smaller models for simpler steps, and hard budget limits. Self-hosted shifts that to fixed hardware cost. Either way I would give you a projected monthly figure before building, not after.

Have a project like this?

Tell me what the system has to do and what it has to talk to. You get a straight answer about scope, sequence and what I would build first — before any commitment.

Start a conversation