About SynthVision

It started with a project that couldn't get its data.

Not a market thesis. A computer vision system was ready to ship and the data it needed was always months away — repeatedly, every time the product changed.

The origin

A vending machine that couldn't see its own products.

The project was a smart refrigeration system: recognising which products were present on each shelf, in real time, from the camera inside the cabinet. The hardware was built. The model architecture was sound. The problem was that the model had nothing to learn from.

Every product change invalidated the dataset. Field teams had to photograph the new SKUs across locations and lighting conditions, at different shelf fills. Then the images had to be reviewed, labelled and validated before training could begin at all.

A catalogue update — something that happened several times a year — took months from first capture to a deployable model. Across a large fleet of units, that lag was not a nuisance. It was the thing blocking the rollout.

"If I can rebuild this shelf in 3D and swap the products parametrically, I can generate the whole dataset in an afternoon."

That was the insight, and it held up. The first synthetic dataset matched what weeks of field capture would have produced, and the next product update took a fraction of the time.

SynthVision exists to make that available to any team whose model is limited by data rather than by architecture — agro, industry, safety, retail, logistics, automotive, robotics and research.

Guilherme Bileki
Founder · Computer Vision

Guilherme Bileki

Computer vision engineer with hands-on experience building production CV systems across agro, retail and industrial domains — from smart devices to full machine learning pipelines.

He founded SynthVision after repeatedly hitting the same wall: getting the right training data takes longer than building the model. Every project is led end to end from diagnosis to iteration, which is also why you're talking to the person building your pipeline rather than a support queue.

Direct access

Sophisticated technology, without the distance of a large supplier.

Large platforms have scale, ecosystem and enormous engineering resources. What they also have is a boundary: if your requirement isn't on the roadmap, your path to it runs through a support queue, a product team and a release cycle.

Enterprise-grade technology.
Direct access to the people who build it.

Here, a requirement becomes an engineering change becomes a new dataset. That's the whole loop — and it's short on purpose.

You talk to the engineer Not an account manager relaying messages to an engineering team you'll never meet.
The pipeline is ours to change New camera distribution, new physics, new class definition, new annotation schema — these are changes, not feature requests.
Small enough to move quickly Speed here comes from proximity, not from headcount. It's the advantage a small team has, and we lean on it deliberately.
We stay after delivery Model performance in development and in production are different problems. The iteration loop is part of the work.
How we think

Positions that shape every project.

Some of these make our job harder. They're still the ones we hold.

The data bottleneck is a model problem

When a model fails in production, the root cause is usually a gap in the data — a condition never represented, a class never balanced, a label that was never exact. Solving that is solving the model.

Design-controlled beats generated

Sampling from a model gives you an image you didn't specify. Declaring a scene gives you an image whose contents and labels you can describe exactly. That difference is the entire product.

Synthetic and real data are complementary

Not a replacement story. Synthetic data extends coverage into conditions real collection can't reach; real data keeps the model grounded and remains the only honest validation set.

Control and traceability are non-negotiable

Every sample traces back to the parameters that produced it. Reproducible, auditable, and explainable — not approximately, and not stochastically.

A technical partner, not a data vendor

Handing over a spec and returning a zip file doesn't work, because the right next batch depends on what the model is still getting wrong. The loop has to stay closed.

Iteration speed compounds

One day from model evaluation to the next training batch, instead of two weeks, is roughly fifteen times the iteration cycles in the same project window. That compounding is where the value concentrates.

Labelling gets easier. Capture doesn't.

General-purpose models keep getting better at annotating images that already exist, and we expect that trend to continue — it's real progress, and it keeps shrinking half of this problem. It doesn't touch the other half: a condition nobody has photographed yet still isn't in anyone's dataset, however good the labelling tool gets. That's the half we build for, and it's not a function of model size.

Want to know how this applies to your project?

We start with a diagnosis call. No deck, no commitment — a technical conversation about your data problem.