A rising share of what gets published on the open web is written by machines, not people.1
1. Measured by multiple independent web content studies. The direction is not in dispute.
Made, not found.
Ultro Labs runs experiments inside your company to capture the record before it fades: clean, human, multimodal data on how the work actually gets done.

The supply problem
The open web keeps growing, but the pool of attributable, human-made source material does not grow with it. Three facts explain why clean records now have to be made at the source.
Continue / three signals
A rising share of what gets published on the open web is written by machines, not people.1
1. Measured by multiple independent web content studies. The direction is not in dispute.
Models trained on model output collapse. The rare cases go first, then the specifics.2
2. Shumailov et al., AI models collapse when trained on recursively generated data, Nature, 2024. Source
The pool of verifiably pre 2022 human writing does not grow. It only gets spent.3
3. A supply constraint, not a forecast. Nothing new is being added to that pool.
The decay
Generation 1 of 8
Generation 1. Marisol has run the returns desk for eleven years. She can tell from the tape on a box whether the customer repacked it at home or in the parking lot, and she knows the difference matters: parking lot returns are impulse regret, home returns are real defects. None of this is in the manual.

From noise to evidence
Every record is captured at the source, tied to the person and context behind it, then tested against a measurable outcome.
What we do
We instrument the work your own people already do, and record the reason behind each call.
A floor camera catching how a technician actually clears a jam, not the order the manual lists.

The same protocol run at several sites, so a finding holds up outside one building.
Four depots running one intake script, so a pattern is a pattern and not a local habit.

Trained collectors capture work where it happens. Multimodal from the first minute.
A collector riding a service route for a week, recording the call and the reason given for it.

Lab layer
Every dataset ships with its own eval.
Held-out cases turn a collection into a testable instrument, measured against your own ground truth.
Specimen library
The asset
A model license expires. A dataset does not. Data collected on your own ground stays on the balance sheet, it is used again with every model you try, and it gets more valuable as the open supply gets worse. You are not renting an answer. You are buying an asset, and you own it outright.
Evals
If it is not measured on your own ground truth, it is a guess.
From the manifesto
The open web is filling with text that no person wrote and no person checked. It is cheap to make and it reads well enough to pass. Every month there is more of it, and every month the average page is a little further from anyone who did the work it describes.
Models trained on that output drift. When a system learns from its own reflections, the rare cases go first, then the specifics, and what is left is a confident average. The finding has a name and a citation: Shumailov and colleagues published it in Nature in 2024. The practical version is simpler. Copies of copies fade.
The work leaves a trace. Enter before it disappears.