Expert data, built around your model
Custom training data, human feedback with rubrics, and evaluations built for your workflows. Every delivery ships with the expert's work, the criteria it was accepted against, and the review decision, so you can check any row.

Custom training data
Expert-written examples, demonstrations, and labels for the tasks your model needs to learn.
See a training recordHuman feedback and rubrics
Expert ratings, response comparisons, and scoring rubrics that turn judgment into consistent feedback.
See a preference labelCustom evaluations
Tests built for your workflows, with expert scoring that reveals failures and measures progress.
See an evaluation reportWhat a delivery looks like
Every record ships with the expert's work, the criteria it was accepted against, and the review decision, so a client can check any row.
Illustrative samples. Deliverables are scoped to your project.A software company invoices a customer $120,000 on 1 January for a 12-month subscription and a one-time onboarding service delivered in the first month. How should it recognise the revenue, and what changes if onboarding is not sold separately?
Two performance obligations. The subscription is satisfied over time, so its share is recognised evenly across the twelve months. Onboarding is satisfied at a point in time, in January, if it is distinct.
Allocate the $120,000 on relative standalone selling prices, not on the invoice split. If onboarding is never sold separately and the customer cannot benefit from it on its own, it is not distinct: fold it into the subscription and recognise everything over the twelve months.
- Identifies both performance obligations
- Allocates on standalone selling price, not invoice
- States the distinctness test and its consequence
- Reasoning trace, 6 steps, cited to ASC 606
- Review record, 2 reviewers, agreed
- Dataset version and rubric v3
A patient on warfarin asks whether they can take ibuprofen for a headache.
| Rubric criterion | A | B | Expert note |
|---|---|---|---|
| Safety: names the interaction | 1 / 5 | 5 / 5 | A omits the warfarin interaction entirely. |
| Accuracy of dosing or alternative | 3 / 5 | 5 / 5 | B gives the standard alternative and defers correctly. |
| Refers to a clinician when warranted | 1 / 5 | 5 / 5 | Required by clause 4.2 of the rubric. |
A is fluent and confident, which is exactly why it fails: it answers a different question than the one that matters for this patient. B is graded on clause 4.2, which the rubric rewrote in v2 after disagreement on the first batch.
| Behavior | Pass | Items | Most common failure |
|---|---|---|---|
| Identifies the governing-law clause | 96% | 60 | Confuses venue with governing law |
| Flags an unbounded indemnity | 81% | 60 | Misses caps expressed by cross-reference |
| Extracts termination notice period | 88% | 60 | Reports days when the clause says business days |
| Declines to answer on missing clause | 54% | 60 | Invents a plausible clause instead of saying none |
Overall accuracy is 80%. The one behavior that matters most to counsel, refusing to fabricate a clause, is the one failing. Every failed item links to the contract, the model output, and the expert's note.
- 240 scored items with expert notes
- Rubric v2 and the v1 diff
- Rerun instructions and dataset digest

No single check is enough. Six of them, composed, are
Every quality method finds one class of defect and misses another. We run the layers your dataset needs and report what each one found. Open a layer to see both.