We removed the thing the company was known for.
Avataar was a 3D augmented reality company. Velocity, its flagship video product, doesn't use 3D at all.
Illustrative. The measured figure is under seven minutes.
The problem
Most products had no video, and the reason was arithmetic.
Studio production ran $500 to $2,000 per video, two to three weeks per SKU. At that price a brand can't keep a current catalogue across one market, let alone twenty-eight. So most SKUs simply went without.
HP made the ceiling concrete. 2,500 product videos across 28 markets, quoted by an agency at $1.25M and six months.
The unit cost had to land two orders of magnitude below studio pricing, or there was no pitch.

The input
A link. That's the whole brief.
Paste a product page. Velocity pulls the images and the brand kit. Picks a layout. Writes and voices the script. Renders it. What comes back is an advertisement ready to run on Meta, not a file in a folder.
The users were brand managers, not editors. That constraint shaped more of this product than the generation quality did.
Brand managers needed a tool that felt like Canva, not like Blender.
The decision
A 3D company shipped its flagship product in 2D.
Avataar built in-room product visualisation. Augmented reality. And AR wasn't being adopted, for concrete reasons rather than vague ones. Phones overheated under sustained rendering. People placed models badly enough to break the tracking.
The company pivoted to images and video. Dropping 3D inside Velocity carried that pivot through to the flagship. Against the organisation's own centre of gravity.
What replaced it: the product is cut out of its photograph. Then placed as a flat image inside a bounding box, in a pre-built template. No model, no scene, no physics.
Scope reduction is the hardest decision to get credit for, because the result is something that doesn't exist.
It was informed rather than brave. They had already watched AR fail on thermals and placement. Removing it was a conclusion, not a gamble.
And it's the decision that makes seven minutes possible. Everything downstream is cheap because nothing downstream is 3D.
How it works
Seven stages, and the interesting one is missing.
Product URL in, runnable advertisement out. The stage where a 3D company would have put a 3D scene is a rectangle and a flat image.
Ingest
The brand manager has typed almost nothing.
Segment
Isolate the product from whatever it was shot against. This is where the product breaks.
Match a template
Not free-form generation. The corpus is annotated by category and subcategory, so it can be selected from.
Place in 2D
A flat image dropped into a rectangle. Where the 3D would have gone.
Voice and localise
28 markets means voiceover and copy per market. Invisible in the headline number, expensive in practice.
Render in parallel
Hundreds at once, on infrastructure that already existed.
Publish
Out as a runnable ad, not an export. The product sits in the campaign workflow, not before it.
Scroll the pipeline sideways →
Why nobody could copy it quickly
The template corpus was built by somebody else's users.
Templates are predictable in a way generation isn't, and predictable output was the binding constraint. But you can't match against a library you don't have.
Creator was a full browser-based video editor. Its 45,000 organic users built templates in it while solving their own problems. Those templates became the corpus.
Then somebody annotated them, by product category and subcategory. That second step is the one people skip when they retell this. It's the step that made matching possible. An unlabelled pile of templates can't be selected from algorithmically.
A competitor needed two things. The templates, which took a product with 45,000 users to generate. And the annotation over them, which is deliberate human labour.
Starting fresh, this choice was simply not available.
The quality bar
Automated, and not obviously automated.
The output had to survive being put on a product page next to studio work. That's the bar an automated pipeline usually fails, and it's why the pipeline matches templates rather than generating freely.
It was never given a measured definition. "Broadcast quality" was a judgement call made by looking, which is honest to admit and a real gap.

The sequence
Thirty days, then fourteen days, then seven minutes.
Velocity is the third act, and it only works because of the two before it. These aren't three separate jobs.
Super Nova
Consolidated 10 to 15 external vendor teams onto one platform. A 30-day content cycle halved to 14.
Creator
A browser-based editor. 45,000 organic users turned their own editing into a template corpus, which is the asset Velocity runs on.
Velocity
Weighted matching over that annotated corpus, with the 3D removed. Product URL to runnable ad.
Roughly a six-thousand-fold reduction, across three products, over five years. Each stage produced exactly what the next one needed.
Where it breaks
Shiny things.
The anchor engagement was laptops.
Automated segmentation fails on reflective and transparent surfaces. The background either shows through or reflects off the thing you're isolating, so there's no correct edge to find.
And this one is worse here than it would be elsewhere. In a normal video tool a bad cut-out gets fixed by an editor. Canva-not-Blender deliberately gave that up, so a bad cut-out has nowhere to go.
The limitation is the cost of the central decision, not a bug in it.
Which is why the fix isn't better segmentation. That's a research problem. Detecting the failure is a product problem. Score the cut-out. Route a low score to a template that doesn't need a clean cut. An inset layout that keeps the original photograph produces a good video from a bad matte. The escape hatch becomes a different template rather than a manual tool, which keeps the design principle intact.
What was never measured
The part that makes the rest believable.
If segmentation fails on glossy surfaces and the anchor engagement was laptops, some products were excluded or handled differently. That isn't written down anywhere.
A pipeline that handles most of a catalogue and says so beats one that claims all of it and quietly degrades. Coverage by category, reported next to throughput, is the number that was missing.
Video beat static imagery in a live test. The control, the sample size, the duration, whether it compared the same SKUs. None of it is recorded. So the number isn't one to lead with.
It was the binding constraint on the whole design and it was assessed by looking at the output. No rubric, no panel, no score.
The pipeline had been fitted to the first client's SKU structure and ERP format. That's the honest cost of an anchor engagement, and it's the thing a second customer always exposes.
