How we work

What we take on, and how we measure whether it worked.

We don't publish client names, client data, or project-specific numbers — not because we don't have them, but because the confidentiality we promise a client covers their results too, not just their identity. What follows is the shape of the problems we take on, in general terms, and exactly how we measure a project without attaching any of it to one engagement.

No client names or data published Confidentiality by default Methodology available on request
Problem patterns

The shape of what we're usually asked to solve.

Described generally, across engagements — not tied to any single client or project.

Retail · industrial packaging · catalogue-driven products

A product line that changes faster than a real dataset can be rebuilt.

The shape of this problem: a catalogue, SKU set or part list that updates often enough that a real-photo dataset is stale before the model is even retrained on it. Field capture, review and labelling take weeks to months per refresh — and the process repeats from scratch on the next change.

How we approach it
  • Rebuild the physical scene as a parametric 3D environment — shelves, packaging, camera geometry
  • Model each product or part as a swappable asset, with the label as an independent property of the asset
  • Generate detection or recognition datasets across the combinations that matter — lighting, occlusion, placement
  • Re-run the same pipeline on the next catalogue change instead of re-capturing from scratch
Why simulation fits

Once the parametric scene exists, a catalogue update is a data-entry change, not a field campaign. The dataset regenerates in the time it takes to render, not the time it takes to schedule a shoot, label it, and validate it.

Agro · industrial defects · safety-critical detection

A condition too rare, or too dangerous, to have been photographed yet.

The shape of this problem: the case that matters — a disease at a specific stage, a defect that occurs once in thousands of units, a near-miss safety event — hasn't shown up in real footage in enough volume to train on, and may never show up on its own schedule.

How we approach it
  • Build the object or scene parametrically, with the rare condition as a controllable variable
  • Vary severity, scale, distribution and surrounding context independently
  • Generate the condition deliberately, at whatever volume the training plan needs
  • Validate against whatever real examples do exist, to check the synthetic distribution is a fair stand-in
Why simulation fits

No amount of faster labelling produces a photo of a condition nobody has recorded yet. This is the case where simulation isn't a shortcut around collection — it's the only source of the example at all.

🔒

We don't publish client names, project-specific data, or outcomes tied to a single engagement — the same confidentiality we'd promise you covers the last client too. If you want to talk through fit for your specific case, including what's realistic to expect, that's a conversation we can have directly.

How we measure

The validation set is the one thing we don't simulate.

Numbers only mean something if you know what they were measured against. Ours are simple to state.

01

Validation stays real

Where real labelled data exists, that is what we report against. Reporting F1 measured on synthetic images would tell you nothing about production behaviour, so we don't do it.

02

The same set for every strategy

When we compare training strategies — real-only, synthetic-only, mixed — they are evaluated on an identical real validation set. Otherwise the comparison is decoration.

03

We'd rather undersell than oversell

If your own real data already gets you most of the way there, we'll say so — even when a bigger synthetic engagement would be the easier sell. The right amount of synthetic data is sometimes none.

04

The method doesn't care which model you use

Same-set validation and controlled test conditions work the same whether the model behind the number is one we helped train from scratch or a fine-tuned foundation model. The measurement approach isn't tied to any one way of building the model.

Working on something similar?

Tell us your domain and where the model currently breaks down. We'll say whether simulation is worth testing on it.