* One-fourth cost is the GoML distilled-SLM benchmark threshold, not the external synthesis result.
Do not rent the frontier model forever. Distil it into an SLM your enterprise owns.
The frontier model is useful for discovering capability. The production model should be designed around the enterprise workload.
A general-purpose model such as Fable 5 carries the cost of being capable across thousands of unrelated tasks. Most enterprise systems do not need that breadth on every request. They need exceptional performance on a bounded set of workflows: understanding internal documents, applying domain policy, producing a fixed output structure, selecting approved tools and escalating safely when evidence is insufficient.
GoML first creates a high-quality teacher system through model synthesis. Multiple models solve, verify and critique the same task; a synthesiser converts their strongest work into a single approved response. GoML then captures those responses, preferences, corrections and tool decisions as training data.
The final step is the commercial differentiator: that intelligence is distilled into one specialised Small Language Model (SLM), optimised and deployed inside the customer environment. The production request no longer needs to call an expensive frontier model or a complete model panel.
Synthesis creates the quality ceiling. Distillation captures the behaviour. SLM creation turns it into your own sovereign asset.
The framework isn't new. Model ensembles, synthetic-data generation, curation, knowledge distillation, fine-tuning and optimised deployment are all established AI techniques. GoML's differentiation is combining them into a governed enterprise model-engineering framework built around the enterprise's own workload and data, measurable acceptance criteria, teacher and student model selection, human and automated data-quality controls, on-premises infrastructure constraints, and cost, latency, quality and privacy targets.
Synthesise
Use complementary frontier teacher models to produce a response stronger than a single-model answer.
Curate
Rank outputs, verify evidence, preserve corrections and build a governed golden dataset.
Distil
Transfer the required task behaviour into the smallest student that clears the benchmark.
Deploy
Quantise and serve the distilled SLM on premise, in a private VPC or at the edge.
Higher task quality at one-fourth the production cost
GoML separates the expensive process of creating intelligence from the economical process of serving it.
Build with a panel. Serve with one distilled SLM.
The synthesis panel is used as a teacher during model preparation and selective high-value escalation. Routine production traffic is handled by one specialised, optimised model.
The four-stage economics
Establish the frontier baseline
Score Fable 5 on the enterprise task suite using identical context, tools, rubrics and output constraints.
Create the teacher ceiling
Run a selected model panel, synthesise the best answer and capture the evidence behind each accepted decision.
Distil for production
Train one smaller model to reproduce the approved behaviour without paying for the panel on every request.
Optimise the runtime
Apply quantisation, continuous batching, KV-cache management and hardware-aware serving inside the customer boundary.
Why the cost drops after distillation
One inference path
A single student model replaces parallel calls to several teachers and a separate synthesiser.
Smaller model footprint
The student can be quantised to a lower-precision runtime selected for the target accelerator and memory envelope.
Higher throughput
Continuous batching, prefix reuse and KV-cache optimisation increase useful work per GPU-hour.
Owned infrastructure
At stable volume, reserved or on-prem capacity replaces variable frontier API charges.
The cited $1.59 and $0.83 figures are external DigitalOcean measurements. One-fourth of the cited Fable 5 baseline is $0.3975, rounded to a ≤ $0.40 per successful task GoML production threshold. The threshold must be measured on the customer's workload after model distillation and infrastructure optimisation; it is not inferred from DRACO alone.
How GoML validates that the distilled SLM is a good alternative to Fable 5
A credible comparison uses the same enterprise inputs, retrieval corpus, tools, context window, output constraints and scoring rubric for Fable 5, the synthesis teacher and the distilled SLM.
| Dimension | How it is measured | Acceptance logic |
|---|---|---|
| Task quality | Weighted, use-case-specific rubric pass rate | SLM must clear the agreed frontier baseline |
| Groundedness | Supported claim ratio and citation verification | No regression on high-risk factual criteria |
| Structured reliability | Schema validation and deterministic checks | Meet production error-rate threshold |
| Latency | P50, P95 and P99 end-to-end response time | Meet interactive or batch SLA |
| Throughput | Requests or tokens per second at target concurrency | Sustain expected workload with headroom |
| Economics | Infrastructure amortisation + energy + operations | ≤ one-fourth of the normalised Fable 5 cost per successful task |
| Privacy | Data-path and network-boundary verification | Sensitive context remains in approved environment |
| Safety and policy | Adversarial and boundary test pass rate | No unacceptable regression in critical controls |
Four Levels of Comparison
Single frontier baseline
Fable 5 and other selected frontier models under the same context and tool conditions.
Synthesis teacher
The best panel configuration, used to establish the attainable quality ceiling.
Distilled SLM
The compact candidate tested for quality retention, latency and cost advantages.
Production shadow
Real traffic comparison before controlled rollout, without affecting user decisions.
Synthesis showed that model architecture can beat a stronger single model
Before distillation, the external benchmark demonstrates the central architectural idea: a carefully selected open-model panel can produce a better answer than Fable 5 while costing less. GoML uses that higher-quality output as teacher data, not as the final high-volume serving architecture.
GoML ran the full DRACO panel for this benchmark, all 100 tasks across all 10 domains, with every individual model and synthesis configuration scored under one common evaluation framework.
GLM 5.2 + Kimi K2.6, synthesised by GLM 5.2
The configuration achieved a 65.65% quality score at $0.83 per task. Fable 5* scored 62.21% at $1.59 per task in the same external test.
This 65.65% is the teacher-stage synthesis result on DRACO, not the distilled SLM's score. The SLM is trained on a task-specific dataset and proven separately on a held-out enterprise benchmark. See the data and training methodology and the test and training infrastructure.
Quality ranking across all 15 configurations
Higher is better. Quality values reproduce the cited DigitalOcean DRACO test; three matrix-only entries did not have costs listed in the published ranking table.
| # | Configuration | Synthesiser | Quality | Cost/task |
|---|---|---|---|---|
| 1 | Fable 5* + GPT-5.6 frontier panel | Fable 5* | 69.01% | $4.76 |
| 2 | GLM 5.2 + Kimi K2.6 value pick | GLM 5.2 | 65.65% | $0.83 |
| 3 | GLM 5.2 + DeepSeek V4 Pro + Kimi K2.6 | GLM 5.2 | 65.21% | $1.23 |
| 4 | DeepSeek V4 Pro + Kimi K2.6 | GLM 5.2 | 64.67% | $0.95 |
| 5 | GPT-5.6 single | — | 64.32% | $1.39 |
| 6 | GLM 5.2 + DeepSeek V4 Pro | GLM 5.2 | 63.87% | $0.99 |
| 7 | GLM 5.2 + DeepSeek V4 Pro | DeepSeek V4 Pro | 62.72% | $0.85 |
| 8 | DeepSeek V4 Pro + Kimi K2.6 | DeepSeek V4 Pro | 62.72% | Not listed |
| 9 | GLM 5.2 + Kimi K2.6 | DeepSeek V4 Pro | 62.38% | Not listed |
| 10 | GLM 5.2 + Kimi K2.6 | Kimi K2.6 | 62.35% | Not listed |
| 11 | Fable 5* single | — | 62.21% | $1.59 |
| 12 | GLM 5.2 + DeepSeek V4 Pro | Kimi K2.6 | 60.83% | $0.89 |
| 13 | DeepSeek V4 Pro + Kimi K2.6 | Kimi K2.6 | 60.33% | $0.83 |
| 14 | DeepSeek V4 Pro single | — | 58.15% | $0.31 |
| 15 | GLM 5.2 single | — | 58.12% | $0.25 |
*Comparability note: DigitalOcean reports that Fable 5 single and synthesis results cover 93 tasks because guardrails blocked seven inputs.
Selected quality comparison
The open two-model value configuration exceeded every single model tested.
The synthesiser matrix reveals the real driver
The panel alone did not determine the result. The same panel could move by several points depending on which model performed the synthesis step.
GLM 5.2
DeepSeek V4 Pro
Kimi K2.6
Synthesiser effect
DigitalOcean reported that changing only the synthesiser shifted panel quality by roughly three to five points.
Adding a third open model
The best two-model panel scored 65.65 versus 65.21 for the tested three-model panel.
Ideal quadrant
Four open combinations reportedly beat Fable 5 on both quality and cost.
Frontier panel premium
The top-quality frontier panel cost $4.76/task versus $0.83/task for the open value configuration.
The benchmark is the starting point; distillation into an SLM adds real value
Synthesis improves the probability that the final response captures more of the rubric. One model may discover a source another misses; another may structure the analysis better; the synthesiser can reconcile overlap, remove weak claims and assemble a more complete answer.
Independent panel responses explore different retrieval paths and reduce dependence on one model's search trajectory.
The synthesis layer can compare contradictory claims, privilege stronger evidence and correct incomplete outputs.
A panel can intentionally combine a strong researcher, a strong domain reasoner and a strong structured-output model.
Long-form tasks reward breadth. Synthesising several candidate reports can increase the number of supported criteria covered.
The benchmark supports a claim about architecture: a well-chosen open-model panel can outperform a frontier single model on a defined task suite.
The external result establishes the teacher-stage quality ceiling. GoML's direct production claim is validated separately: the distilled SLM must beat the agreed Fable 5 baseline on the enterprise benchmark while meeting the ≤ $0.40 cost-per-successful-task threshold.
How the supporting research benchmark was constructed
DRACO measures Deep Research Accuracy, Completeness, and Objectivity. It is useful here because it tests long-form, evidence-heavy research rather than one-shot recall. It remains supporting evidence, not the centre of GoML's SLM proposition.
DRACO is designed for long-form, evidence-heavy research agents rather than one-shot question answering. Its 100 tasks were derived from anonymised real-world deep-research usage and span ten domains with source requirements covering 40 countries.
Two separate datasets are in play here, and they should not be conflated. For the external model-synthesis benchmark, DigitalOcean used the public DRACO dataset to compare four individual models and 11 synthesis configurations. DRACO was not used to train a GoML distilled SLM. For GoML's own enterprise distillation process, the training dataset is built from a customer's specific workload: existing enterprise data, historical production requests, approved outputs, subject-matter-expert examples, policy and workflow documentation, synthetic examples generated by teacher models, difficult and negative cases, and human-reviewed golden responses.
DRACO was selected for the external synthesis experiment because synthesis is particularly relevant to complex questions that require research across multiple sources, reconciliation of conflicting information and detailed, cited answers. DigitalOcean chose it because its breadth and evidence requirements align with what model synthesis is intended to improve. An enterprise benchmark, by contrast, is chosen for production relevance: task frequency, business value of a correct response, risk of an incorrect response, coverage of common and uncommon cases, representative terminology and formats, availability of reliable expected answers, inclusion of adversarial and out-of-scope inputs, and a strict separation of training, validation and test examples to prevent test-data leakage. This benchmark is defined before training and evaluated against a fixed, unseen test set.
In short, there are two separate proofs. DRACO shows that model synthesis can outperform Fable 5 on complex deep-research tasks. A separate enterprise held-out benchmark then shows whether the distilled SLM preserves or improves that performance on the specific workload it was trained for.
Complex tasks
Open-ended questions requiring multi-hop retrieval, cross-source reasoning and a cited report.
Knowledge domains
Finance, product comparison, academic, technology, general knowledge, UX, law, medicine, needle-in-a-haystack and personal assistant.
Domain experts
Medical professionals, attorneys, analysts, engineers and designers helped create and validate the rubrics.
Criteria per task
The dataset contains 3,934 weighted criteria, including 415 negative criteria that penalise serious errors.
How the rubric weight is distributed
How answers are graded
Each answer is checked criterion by criterion using an LLM-as-judge protocol. The judge produces a binary verdict for each criterion, then the weighted criterion results become the task score. The DRACO authors tested multiple judge models: absolute scores varied, but relative system rankings remained consistent.
DRACO evaluates deep-research systems with browser and code-execution capabilities. It does not prove that one model or panel is universally superior for coding, extraction, customer support, clinical coding or other enterprise tasks. It is strong evidence for this specific class of research workflow.
Additional reading
- Official DRACO research paper — methodology, task construction and evaluation design.
- Public DRACO dataset — tasks and associated evaluation rubrics.
- Perplexity research explanation — an accessible overview of how the benchmark was developed.
- DigitalOcean model-synthesis experiment — source of the Fable 5 comparison used in this article.
How GoML turns model synthesis into a production SLM
The GoML lifecycle is designed backwards from the production outcome: one specialised model that preserves the required quality, meets the cost threshold and operates inside the enterprise security boundary.
GoML Synthesis → Distillation → Deployment pipeline
A governed path from enterprise tasks to private model weights.
Domain benchmark
Define tasks, rubrics, baselines, risk cases and acceptance thresholds.
Teacher panel
Select complementary frontier and open models by task capability.
Evidence synthesis
Run parallel answers, critique, verify and consolidate against rubrics.
Gold dataset
Create ranked outputs, preference pairs, hard negatives and tool traces.
SLM distillation
Fine-tune the smallest viable student and preserve required behaviour.
Private runtime
Quantise, package, secure, monitor and deploy on target infrastructure.
1. Build the domain benchmark before building the model
The evaluation dataset becomes the design contract. GoML starts with representative production questions, difficult edge cases, policy-sensitive scenarios, structured-output requirements and explicit failure conditions.
2. Select teachers by capability, not popularity
The optimal panel is use-case dependent. A coding workload may combine different teachers from a regulatory-research workload. GoML evaluates each teacher on the enterprise benchmark and assigns roles such as researcher, verifier, planner, formatter or critic.
Teacher selection is an evidence-led pilot, not a fixed pairing that is assumed to work for every enterprise. In DigitalOcean's external test, the choice of synthesiser alone shifted panel quality by roughly three to five points, and the strongest two-model panel slightly outperformed the tested three-model panel — evidence that model role and compatibility matter more than simply adding more models. Teacher models are selected on:
- Baseline quality on the enterprise dataset
- Complementary capabilities
- Domain knowledge
- Reasoning and structured-output quality
- Ability to critique other responses
- Cost and latency
- Context-window requirements
- Data-handling terms and licensing
- Availability in the target deployment region
3. Generate more than final answers
Useful distillation data includes accepted answers, rejected answers, critiques, corrected responses, evidence maps, citations, tool-selection labels, abstention examples and policy boundaries. This gives the student a richer learning signal than a flat instruction-response dataset.
AI-ready knowledge
Cleaning, deduplication, PII masking, ontology alignment and train/eval separation.
Teacher intelligence
Parallel generation, rubric scoring, evidence checking and response reconciliation.
Specialised student
Supervised tuning, preference optimisation and targeted corrective training.
Production lifecycle
Canary release, drift monitoring, regression suites and continuous refresh.
Distil the exact behaviour that made the teacher system better
Distillation is not generic compression. GoML transfers the specific reasoning patterns, corrections, preferences, formats and tool decisions that matter for the customer workflow.
| Distillation layer | Training signal | What the student acquires |
|---|---|---|
| Response distillation | High-scoring final answers | Domain answer patterns and expected completeness |
| Preference distillation | Chosen vs rejected outputs | Organisation-specific style, accuracy and decision preferences |
| Corrective distillation | Critique → rewrite pairs | Known failure avoidance and self-correction patterns |
| Tool-use distillation | Retrieve / call / escalate labels | Reliable interaction with enterprise systems and humans |
| Boundary distillation | Abstain, refuse and escalate examples | Safe operation inside defined policy and evidence limits |
| Format distillation | Schemas and validation results | Stable JSON, code, clinical or operational output contracts |
The objective is not "a smaller Fable 5." The objective is the strongest model for one enterprise operating envelope.
Distillation itself is not a new idea. Model compression techniques were explored as early as 2006, and modern knowledge distillation was formalised by Hinton, Vinyals and Dean in 2015. More recent research has extended distillation to LLM-generated explanations and reasoning supervision. This is the approach GoML applied when building enterprise teacher panels.
Choosing the student model
GoML benchmarks several model sizes rather than assuming that a larger student is always better. The chosen model is the smallest configuration that clears the required quality threshold while satisfying latency, memory, throughput, licensing and deployment constraints.
Student models are selected on:
- Performance before fine-tuning
- Required model size
- Target GPU, CPU and memory constraints
- On-premises deployment compatibility
- Context length
- Fine-tuning and quantisation support
- Commercial licence
- Target latency and throughput
- Language and domain coverage
The final teacher-and-student combination is determined by piloting on the enterprise's own workload, not by picking a single pairing off a public leaderboard.