Service
Agent & RAG evaluation
A scoped investigation that tells a Head of AI or release owner whether an agent or retrieval change is ready to ship, what it breaks, and why, using a task suite built from your own workflows and a regression harness you keep.
The decision
Can we release this agent or retrieval change, and what will break when we do?
Usually commissioned by
Head of AI · AI platform lead · Release owner · Search and retrieval lead
You receive
- Task-based evaluation suite built from your real workflows, with graded reference outcomes
- Regression harness that runs in your environment and stays with your team
- Failure taxonomy: where the agent or pipeline fails, how often, and on which task types
- Retrieval diagnosis separating relevance, grounding and generation failures
- Release recommendation against thresholds agreed before any run
The decision this answers
Someone in your organisation has to sign off a release. It might be a new agent workflow, a model swap underneath an existing assistant, a re-chunked document index, or a change to the tools an agent may call. The demo went well. The question on the release owner’s desk is narrower and harder: on the work this system will actually be asked to do, does the new version succeed often enough, and when it fails, does it fail in ways you can live with?
That question is hard for three reasons. First, agents and retrieval pipelines fail on the long tail of real requests, not the handful of examples a team rehearses. Second, a single accuracy number hides which kinds of task got worse; a release can improve the average while breaking the one workflow your largest customer depends on. Third, retrieval systems fail in layers. A wrong answer may come from the retriever returning the wrong passages, from the model ignoring the right passages, or from the model asserting something no passage supports. Each has a different fix, and a score that blends them tells you none of it.
How we build the evaluation
We start from your workflows, not from a public benchmark. With the people who own those workflows, we sample real or realistic requests, write down what a correct outcome looks like for each, and agree how it will be graded. Some tasks have checkable answers: a record updated, a SQL result matching, a citation pointing to the right clause. Others need rubric grading by a person or a model, and where a model grades, we measure its agreement with human graders on a sample before trusting it.
Thresholds are agreed before anything runs. A release recommendation is only meaningful if the bar was set before the result was known, so the plan states, for example, the minimum task success rate on each workflow, the maximum tolerated rate of unsupported claims, and which failure types are disqualifying regardless of the average.
For retrieval systems we instrument each layer separately:
- Retrieval relevance: did the right passages appear in what was returned, and at what rank?
- Grounding: given the passages returned, did the answer stay within what they support?
- Generation: was the answer correct, complete and in the form the workflow needs?
For agents we record the full trajectory, including tool calls, arguments, retries and hand-offs, so a failure can be attributed to a step rather than to “the agent”. Those traces become the failure taxonomy: a named list of failure types, how often each occurs, on which task types, and an example trace for each.
What you receive
- A task suite drawn from your workflows, with reference outcomes and grading rules written down.
- A regression harness that runs in your environment, against your deployment, and belongs to you after the engagement.
- The failure taxonomy, with counts, examples and the slice of tasks where each failure concentrates.
- A retrieval diagnosis, where retrieval is in scope, stating which layer is responsible for which share of failures.
- A release recommendation against the agreed thresholds: release, release with named restrictions, or do not release, with the reproducible evidence for whichever it is.
The report separates what was measured on your tasks from what we infer beyond them and from what remains unknown. A sample decision report shows the format.
The instruments, and when we leave ours out
Skelf maintains open-source tools that cover parts of this work. They are instruments, not a product you must adopt, and the right harness is often your own stack or a third-party evaluation framework. We choose per engagement and explain the choice.
| Instrument | What it can do in an evaluation | When we would not use it |
|---|---|---|
| Promptel | Express prompt variants declaratively with typed parameters across providers | Your prompts already live in a templating system you trust |
| Blogus | Pin each prompt by sha256 in a lock file so the prompt evaluated is provably the prompt shipped | You already version prompts with equivalent integrity |
| MPL | Sit between agents and their tools, validate payloads against typed contracts, score quality and write hashed audit records | Your agent framework already enforces schemas and logs traces you can audit |
| EmbedCache | Fix embeddings by content hash and model id so retrieval reruns are reproducible | Embeddings come from a managed service whose versioning you already control |
| Memista | A small local vector index to reproduce retrieval behaviour outside production | Above roughly 100,000 vectors, or for anything resembling production load |
| Polymathy and Slorg | Web-retrieval pipelines; Slorg’s fixed plan-first pipeline is a useful non-agentic baseline | The system under test does not search the open web |
| l0l1 | Validate agent-generated SQL against the live schema and scan for PII before queries leave | The agent never writes SQL |
Two of these carry boundaries worth stating plainly. MPL’s audit records and quality scores are evidence primitives, not certifications; using it does not make a system compliant with anything. Memista is experimental at v0.1.x and tested below about 100,000 vectors, so we use it only to reproduce small retrieval scenarios, never as a recommendation for your production index.
What we need from you
- Workflow owners for two or three working sessions, to agree tasks, correct outcomes and thresholds.
- Representative requests, real where possible and redacted where necessary. Synthetic tasks fill gaps but are labelled as such in the report.
- Access to the system under test in a staging environment, including the ability to run the current and candidate versions side by side.
- The retrieval corpus or a faithful snapshot of it, where retrieval is in scope, along with the chunking and embedding configuration.
- A named release owner who will receive the recommendation and decide.
Customer data stays under the terms agreed in procurement. Our methods are public; your tasks, traces and results are confidential unless you expressly agree otherwise.
What this will not tell you
It will not tell you that your system is “trustworthy AI” in general. It tells you how a specific version performed on a specific task suite against thresholds you agreed. It does not cover workflows you did not include, adversarial users unless red-teaming is explicitly scoped, or behaviour after the next model or data change. It is not a security audit or a regulatory assessment.
We also have to be clear about our own record. Skelf has not yet published a measured study of agent evaluation. A study comparing single- and multi-agent designs on public tasks is planned; until its results are in the evidence register, the method on this page rests on established evaluation practice and on our published work in other domains, not on a Skelf agent benchmark.
How it typically runs
A technical diagnostic (£2,500–£5,000) frames the release decision, reviews whatever evaluation already exists, samples tasks with your workflow owners, agrees thresholds and produces an experiment plan. Sometimes the diagnostic alone settles the question; it often reveals that the existing test set never exercised the workflow that matters.
An evaluation sprint (£8,000–£20,000) builds the suite and harness, runs current against candidate, writes the failure taxonomy and delivers the release recommendation.
Recurring re-evaluation ties reruns to change events rather than the calendar: a model upgrade, a prompt edit, an index rebuild, a new tool. Because the harness and the pinned prompts already exist, a rerun answers “did this change break anything we care about?” without repeating the sprint. The methods page describes how we record measured, inferred and unknown results across all of these.
| Engagement | You receive | Price (GBP) |
|---|---|---|
| Technical diagnostic | Problem framing, baseline review, experiment plan, scope and recommendation | £2,500–£5,000 |
| Evaluation or benchmark sprint | Controlled comparison, reproducible harness, failure analysis, decision report | £8,000–£20,000 |
| Applied R&D project | A bounded prototype or new method, experimental results, limitations, handover | £25,000–£75,000 |
| Recurring re-evaluation | Agreed testing after model, data, infrastructure or policy changes | £2,000–£8,000 / month |
| Sponsored research package | Defined research outputs with disclosed funding and publication terms | £15,000–£50,000 |
Prices exclude VAT. Compute, external reviewers, travel and third-party licences or data are quoted separately. A reproducible negative finding satisfies the contract; nothing is priced on a favourable result.
Decision guides
Instruments we may use
Chosen only where they suit the question. Maturity and licence are checked per engagement.
Questions buyers ask
- Do we have to adopt Skelf tools to commission an evaluation?
- No. The harness is built on whatever suits your stack, which is often your own tracing and test infrastructure or third-party evaluation tools. We use a Skelf instrument only when it does a specific job better than the alternatives, and we say why in the plan.
- What happens if the change does not meet the threshold?
- Then that is the finding, and it satisfies the engagement. You receive the failing cases, the failure taxonomy and the harness that reproduces them, which is usually more useful than a pass.
- Can you certify our agent as trustworthy or compliant?
- No. We measure a defined system on defined tasks against defined thresholds. That is evidence for a release decision, not a general property of the system and not a certification.
- Has Skelf published its own agent-evaluation study?
- Not yet. A study comparing single- and multi-agent designs on public tasks is planned and will appear in the evidence register when it has results; until then, nothing on this page should be read as a measured claim about agents.
- When should we re-run the evaluation?
- Whenever something the result depends on changes: a model upgrade, a prompt change, an index rebuild, a new tool in the agent's loop, or a shift in the documents being retrieved. The harness is designed so those reruns are cheap.
Is this decision on your desk?
Tell us the decision, the deadline and what evidence would settle it. We reply with whether we can help, and if so the smallest investigation that would.