Sample deliverable
Sample decision report: smaller-model replacement template
This is a template. It shows the sections, tables and sign-offs a Skelf decision report contains, using a fictional team deciding whether a smaller model can replace their current one for document extraction. Every value is a placeholder; nothing here is a measured result.
About this template
This page is an illustrative template, not a report from a real engagement. The scenario is fictional: a team (called “the client” below) deciding whether a smaller language model can replace their current one in a document-extraction workflow. Every value shown is a placeholder. Text in square brackets, such as [result] or [n] tasks, marks where a real report would put a measured figure, a name or a date. No number on this page is a measurement, an estimate or a typical result.
The point is to show a buyer what a decision report contains before they commission one: the order of sections, the tables, and where the sign-offs go. The structure below follows the LLM cost and quality benchmark offer, and the decision it supports is described in replacing a model with a smaller one.
1. Decision and recommendation
Recommendation: [switch / do not switch / switch for a defined subset only].
On the agreed evaluation set, the candidate [candidate: model B] scored [result] against the baseline [baseline: model A] on field-level F1, which is [within / outside] the acceptance threshold of [threshold: ≤ 2 points below baseline on field-level F1]. On the hard slice [slice name] it scored [result]. Cost per accepted result changed by [result], and p95 latency by [result]. On that evidence we recommend [outcome].
“Do not switch” is a legitimate result of this report and is written the same way: the candidate did not meet the agreed threshold, here is where it fell short, and here is the evidence that lets you check it. A recommendation to keep the current model is a finished deliverable, not a failed project.
2. Question, scope and what is out of scope
Question. Can [candidate: model B] replace [baseline: model A] for extracting [fields: e.g. parties, dates, amounts] from [document type] without breaching the acceptance threshold in section 4?
In scope. The extraction step only, on [n] documents drawn from [source and date range], using the client’s current prompt [prompt version / hash] and output schema [schema version].
Out of scope. Upstream OCR quality, downstream human review workflow, fine-tuning either model, and any document type not represented in the evaluation set. The report makes no claim about those.
3. Baseline and candidates
| Role | Model and version | Deployment | Prompt | Notes |
|---|---|---|---|---|
| Baseline | [baseline: model A, version] | [provider / self-hosted] | [prompt hash] | Current production configuration |
| Candidate 1 | [candidate: model B, version] | [provider / self-hosted] | [prompt hash] | Same prompt as baseline |
| Candidate 2 | [candidate: model B, version] | [provider / self-hosted] | [prompt hash] | Prompt adapted for the smaller model |
Including a prompt-adapted candidate separates “the smaller model is worse” from “the prompt was tuned for the larger model”.
4. Acceptance threshold
The threshold was agreed before any measurement and is not revised after results are seen.
- Primary:
[threshold: ≤ 2 points below baseline on field-level F1]on the full set. - Hard slice:
[threshold on slice]on[slice name]. - Hard constraints:
[no regression on fields marked critical], p95 latency[≤ limit]. - Value condition: cost per accepted result
[must fall by at least / any reduction].
| Role | Name | Date signed |
|---|---|---|
| Decision owner | [name, title] | [date] |
| Technical lead | [name, title] | [date] |
| Skelf investigator | [name] | [date] |
5. Method
- Datasets.
[n] tasksfrom[source], frozen at[date]and referenced by[dataset hash]. Labels from[who labelled, how disagreements were resolved]. - Slices.
[slice list: e.g. multi-page documents, scanned inputs, non-English], defined before measurement. - Repetitions.
[n] runsper model per task at[temperature], to measure run-to-run variance. - Grading. Field-level exact and normalised match by
[deterministic scorer, version];[n] disagreementsadjudicated by[role]. - Controls. Baseline re-run against itself, to set the noise floor for the comparison.
- Versions pinned. Model identifiers, provider API version, prompt hashes, scorer version and harness commit
[commit].
6. Results
| Measure | Baseline | Candidate 1 | Candidate 2 | Threshold | Met? |
|---|---|---|---|---|---|
| Field-level F1, overall | [result] | [result] | [result] | [threshold] | [yes / no] |
| Field-level F1, hard slice | [result] | [result] | [result] | [threshold] | [yes / no] |
| Cost per accepted result | [result] | [result] | [result] | [condition] | [yes / no] |
| p95 latency | [result] | [result] | [result] | [limit] | [yes / no] |
Each cell in a real report carries an interval from the repeated runs, not a single point.
7. Failure analysis
| Failure type | Baseline count | Candidate 1 count | Candidate 2 count | Example reference |
|---|---|---|---|---|
| Field missing | [n] | [n] | [n] | [task id] |
| Wrong value, correct field | [n] | [n] | [n] | [task id] |
| Value from wrong document section | [n] | [n] | [n] | [task id] |
| Schema violation | [n] | [n] | [n] | [task id] |
| Hallucinated field | [n] | [n] | [n] | [task id] |
The narrative under this table explains which failure types account for any gap, whether they cluster in a slice, and whether a prompt or post-processing fix is plausible.
8. What was measured, inferred and unknown
Measured
- Accuracy, cost and latency of each model on the frozen set, under the pinned versions.
- Run-to-run variance at the stated temperature.
Inferred
- Behaviour on future documents of the same type, assuming the mix stays similar to
[date range]. - Monthly cost at
[volume], from measured cost per result.
Unknown
- Behaviour on document types not in the set.
- Behaviour of later model versions, which the provider may change without notice.
9. Limitations and when the recommendation expires
Limitations: [sample size caveat], [label quality caveat], [slice coverage caveat].
Re-evaluate when any of these happens: the provider releases or retires a version of either model; the document mix shifts, for example a new supplier or template; the prompt or schema changes; pricing changes; or [date] passes, whichever comes first.
10. Reproducibility package
- Evaluation harness at commit
[commit]. - References to the frozen dataset and labels by hash (the data itself stays with the client).
- Model, prompt and scorer configurations.
- Seeds and run parameters.
- Raw model outputs and per-task scores.
- A
make reproducetarget that re-runs the comparison and regenerates every table in this report.
11. Confidentiality
Client documents and labels remain with the client and are not copied into Skelf systems beyond what the engagement needs. Nothing from this report, including the fact of the engagement, is published without the client’s written agreement.
Related service: LLM cost–quality benchmark · All evidence
Questions buyers ask
- What format is the decision report delivered in?
- A written report, usually as PDF and Markdown, with the reproducibility package alongside it as a repository or archive. The sections follow this template, though the tables change to fit the question.
- Who owns the report and the evidence?
- You do. The report, the raw outputs and the harness configuration for your data are yours. Skelf keeps the right to reuse its own general-purpose methods and open-source tools, but not your data or your results.
- Can we share the report inside our organisation?
- Yes. It is written so that someone who was not in the project, such as a CTO, a risk reviewer or a procurement lead, can follow the decision and check how it was reached.
- What happens if the result is negative?
- A negative result is a valid outcome. If the candidate does not meet the threshold agreed before measurement, the report recommends not switching and includes the reproducible evidence explaining why. The engagement is complete either way.
- How long does the recommendation stay valid?
- Until one of the re-evaluation triggers listed in the report fires, such as a new model version, a change in the document mix or a price change. Recurring re-evaluation re-runs the same harness when that happens.
Is this decision on your desk?
Tell us the decision, the deadline and what evidence would settle it. We reply with whether we can help, and if so the smallest investigation that would.