Model tuning & alignment
Supervised fine-tuning, preference optimization and domain adaptation on open and frontier models — with the eval harness built before the first training run, not after it.
Applied AI research & engineering
Tunedness is an applied AI lab. We take models that almost work — the promising prototype, the eval that plateaued, the agent that breaks in production — and tune them into systems your team can actually operate.
Selected partners
We work across the whole path from research to production, but we are usually called in for one of these. Engagements start with a diagnosis, not a proposal.
Supervised fine-tuning, preference optimization and domain adaptation on open and frontier models — with the eval harness built before the first training run, not after it.
Chunking, indexing and re-ranking designed around how your people actually ask questions — not around a vector database's defaults.
Tool-using systems with explicit state, recovery paths and cost ceilings. Built to be debugged at three in the morning, not demoed once on a good day.
Task-specific eval suites, offline and online, plus the traces to explain a regression. If a change cannot be measured, we do not ship it.
Quantization, distillation, caching and model routing. The goal is the same answer quality at a token bill your finance team stops asking about.
A standing research team for organizations whose problem no off-the-shelf model solves. Scoped in weeks, reviewed in the open, owned by you at the end.
A standing share of our time goes to open problems we keep hitting in the field. Three of them are permanent fixtures.
Notes & preprintsA model tuned on last quarter's data meets this quarter's reality. We study guardrails that degrade honestly — refusing, escalating, flagging low confidence — instead of failing silently and confidently.
Public benchmarks are a starting gun, not a finish line. We build eval methodology whose score still predicts real quality six months after launch, and we publish what stops working.
How small can a model be and still do one job properly? Distillation, curriculum design and task decomposition, aimed squarely at the cost curve rather than the leaderboard.
Every engagement runs on the same internal stack we built for ourselves. When the work ends, the evals, the lineage and the dashboards stay with you — running in your own infrastructure.
Every prompt, weight and dataset change produces a comparable run. Regressions are found by the pipeline, not by a customer.
Trace any production answer back to the checkpoint, the prompt revision and the training slice that produced it.
Latency and cost ceilings enforced at the router. A runaway agent hits a wall before it hits your invoice.
Illustrative interface — values are placeholders
[N]
Systems in production
[N]
Researchers & engineers
[N×]
Median cost reduction
We stay deliberately small so that the person answering your question is the person who ran the experiment. Three rules hold that together.
Weekly written updates, failures included. You see the eval numbers the same day we do — not a polished version of them a month later.
The people on the first call are the people writing the code. There is nobody in between translating your problem into a statement of work.
Weights, prompts, eval sets, infrastructure code, runbooks. We leave holding nothing that you need in order to keep going.
We reproduce your failure ourselves, instrument it, and write down what is actually wrong. The diagnosis is yours whether or not we go further.
A small team working inside your repository, shipping behind a flag, gated by evals from the very first commit rather than the final review.
Runbooks, eval suites and paired sessions with your engineers until they are running it without us. Then we get out of the way.
Start with the hard part
A technical call with the people who would do the work. No deck, no discovery phase — bring the eval scores and the traces you are stuck on.
Book a technical callor write to hello@tunedness.com