Give it your product photos,
get back a full set you can shipImagery, video, 3D and voice — all from one visual standard
In
Photography you already have
Out
Imagery · video · 3D · audio
Runs on
Your own servers
Delivery
In blocks · scoped at the diagnostic
The pipeline in one line
Photography in, a shippable set out
What goes in
The product photography you already have. No reshoot, and no need to organise it into a particular format first.
What comes out
Scene imagery, sized hero and detail images, social crops and short-form video source material — consistent across the batch, with filenames and alt text generated, ready to write back to your catalogue.
The part in between
Runs on your own servers. Material, configuration and models never leave your boundary, and nobody is billing you per image.
What actually goes wrong
It isn't ugly, it's uncontrollable
Different every time
The same prompt ten times gives you ten different looks. Fine for one image, fatal for a product range.
You can't specify it
“Product on the left, space above” — a text description can't control layout precisely. Every revision is another roll of the dice.
The product comes out wrong
A general model has never seen your product, so the details are invented. On a commerce detail page that's disqualifying.
And one more practical one: outsourced product imagery costs tens of currency units per image, takes days, and still doesn't match across a set. Launch often enough and artwork becomes the bottleneck for everything downstream.
Not only imagery
Four kinds of output, one visual standard
Imagery
Scene shots, cut-out heroes, detail images, social formats. Every size exported in one pass with its own safe area, margins and subject proportion — not one crop resized; filenames and alt text generated with them, ready to write back to your catalogue.
A batch arrives as a set
Video
Product motion, scene shorts, vertical social cuts. Built from the same approved images and the same visual standard, so the video and the stills don't contradict each other.
One step after the imagery
3D
A usable three-dimensional model from product photography or a description, then renders from any angle, structural breakdowns and exploded views. Resize it, re-lay-out the breakdown — that part is computed on your own machine, so another version costs nothing extra.
Another version costs nothing
Audio
Voice-over, subtitles, background music. Every language version comes from the same script, so the wording doesn't drift in translation.
One script, every language
You don't need all four. Most engagements start with imagery and extend once that runs — the foundation is shared, so each class after the first is noticeably lighter.
Three tiers
Not one method — a choice between three
Straight generation
A general model generates the image outright. Suited to mood shots, backgrounds and social filler — anything where the product doesn't have to be exactly right. High volume, fast turnaround.
The limit: the product's actual form is guessed. Not acceptable for a hero or detail image.
Constrained generation
Your real photograph is the reference the output is bound to: the product itself isn't repainted, only the background, lighting and setting are rebuilt. Framing and silhouette hold, a batch comes out as a set, and it can go straight onto a product page.
The prerequisite: the product needs one usable photograph to work from.
Brand-specific model
A model trained on your own product photography, so it knows your products specifically. After that you no longer need a reference shot every time to get the form right — including angles and settings photography can't practically reach. Worth it with many SKUs, frequent launches, or heavy demand for scenes you can't shoot.
The model is your asset. It lives on your servers and you can take it with you.
The three mix. Hero images at tier two where accuracy is non-negotiable, scenes and social at tier three for volume, mood material at tier one. We give a recommendation per category during the diagnostic — including “tier two is enough, don't pay for tier three”where that's the honest answer.
How a batch gets through
Four stages, one of them yours
Your own team runs the next launch without us. For scale: selection takes us roughly 20 seconds an image, about eight minutes for a batch of 24 — but that's our own pace on our own catalogue, not a commitment about yours. Your batch sizes, categories and review standards differ, and we estimate against them during the diagnostic.
Video and audio
The same material, extended into video and voice
What's delivered is repeatable production capacity, not a folder of finished clips. Next month's launch runs the same pipeline with your own people. That's the difference from commissioning a batch of videos: that budget is spent and gone, this one leaves something behind.
Not included in this part
- ✕Physical photography
- ✕Voice talent and on-camera models
- ✕Scripting and creative direction
3D
Some products can't be explained by a flat image
Renders from any angle
One model, any viewpoint — including the ones no camera could be rigged for. Changing the angle doesn't mean scheduling another shoot.
Breakdowns and exploded views
How the parts go together and in what order, in one image. Useful for assembly instructions, support tickets and dealer training alike.
Sizes and configurations
Different sizes and specs in the same range: change the parameters and re-output. That part is computed on your own machine, so another version costs nothing extra.
Two boundaries on this one
One: whether it comes apart depends on how the model was made. A model converted straight from a photograph is usually a single shell and won't separate into parts; an exploded view needs a model built structurally. That gets settled at the diagnostic, not discovered before delivery.
Two: this is not engineering drawing. It shows what a thing looks like and how it goes together, not tolerance-accurate parts you could tool from. For that precision, you want a structural engineer.
Brand-specific model
What training one asks of you
Three that don't bend
Each one narrows the tool. Together they make the batch usable.
It has to be that product
If a customer will use the image to judge what the thing actually looks like, it is generated from your own photograph as the source of truth and the product itself is never repainted. We give up freedom for that deliberately — something that merely resembles what you sell comes back to you as a return.
The model never sets type on an image
Let it, and some fraction of the time the text is blurred, misspelled or a letter long — and fixing one word means re-running the whole image. So text is composited afterwards: exact position, correct glyphs, ten language versions of one image in a single pass, and changing a word re-runs nothing.
Spend has a gate, not a warning
Every call is priced before it runs, and one that would exceed the day's ceiling is refused rather than flagged for you to decide about; actual spend is ledgered per call and queryable by section and by day. Video costs one to two orders of magnitude more than an image, so this has to work first.
These came out of the desk we built for ourselves; the full set is on our own visual production desk →
What you end up with
A pipeline that runs itself, not a folder of images
A repeatable pipeline
The same configuration next month produces a set that matches this month's. Nobody has to remember how it was tuned.
Your brand-specific model
Delivered when you go to tier three: trained on your product photography, stored on your servers, dependent on no platform.
An export standard
Sizes, spacing and naming rules for hero, detail and social, written down and handed over so a new hire can follow it.
A run log
Who approved each batch, when, and how many frames came out — traceable, and resumable from where it stopped.
Not included
- ✕Physical photography: materials, real usage settings, models
- ✕Models, locations and location shoots
- ✕Voice talent, on-camera models, scripting and creative direction
- ✕Holding your ad accounts or making budget decisions
- ✕Content compliance and legal review
When it isn't worth doing
- ✕A few launches a year and a few dozen images total — outsourcing is cheaper
- ✕No brand visual standard yet: a machine executes a standard, it can't decide one for you
- ✕Products that look different every batch — there's no stable feature for a model to learn
- ✕You want a handful of striking campaign images — that's creative work, not pipeline work
What this looks like wired into a storefront: E-commerce automation & support agent →
Questions
What people ask about this one
01Which tier do we need?+
One question decides it: will a customer use this image to judge what the thing actually looks like? If yes — hero images, detail pages — start at tier two, where the product's form has to be correct. If no — mood shots, backgrounds, social filler — tier one is enough and anything more is wasted money. Tier three earns its cost when you have many SKUs, launch often, or need volumes of angles and settings photography can't reach. During the diagnostic we give a per-category recommendation, including “tier two is all you need” where that's the answer.
02Can the output go live as-is?+
Yes, but it passes one of your people first. Selection is the step we deliberately don't automate — a hero image moves your return rate, and nothing in that category should publish without a human having looked at it. The machine prepares the candidates and every size; which version ships is your call.
03Do we need a brand visual standard first?+
Ideally yes. If there isn't one we'll set it once and write it down, and it stays fixed after that so a new hire can follow it. But to be clear about the boundary: a machine can execute a standard, it can't decide what your brand should look like. Without that decision, the output can only be attractive — it can't be recognisably yours.
04Is the custom model a recurring cost?+
No. Training material, configuration and the model file are all on your own servers and belong to you; you can keep training new versions or hand it to someone else. We don't bill per image or per call — how much you generate doesn't change what you pay us. The other half should be just as clear: generating images does cost money. That cost isn't in the delivery fee and doesn't pass through us — you top up your own account, the compute runs on your own servers, and we take no cut and add no margin; spend is visible per module, per person, per month. What one image costs depends on the model, the output size and how many passes it takes: on our own line it works out at around $10 a shot, against $30 outsourced (source: image production log). Your categories and sizes differ, so the per-image figure and a realistic monthly total are worked out against your actual volume during the diagnostic.
05Does AI-generated imagery need to be labelled?+
In most markets now yes, and some of it is already in force. China's labelling measures for AI-generated content took effect on 1 September 2025, requiring both visible labelling (on-image marks, text notices) and embedded labelling such as file metadata. The EU AI Act's transparency obligations apply on a timetable from August 2026. Several ad and social platforms have their own disclosure rules, and non-compliance can get an ad rejected or throttled. For this line it comes down to one thing: the label should live inside the pipeline rather than depend on somebody remembering to add it — marks, notices and metadata can all be written in at export, and that part is in scope. What is not in scope: deciding how your market and your channels specifically require you to label. That is a compliance question, the rules differ by region and keep changing, so work from current local requirements and your own counsel.
06Who owns the imagery it produces?+
Positions differ by jurisdiction, so only the broad shape is worth stating: output generated almost entirely by AI, with little human authorship, is generally hard to claim rights over; the human contribution — prompting, selecting, arranging, editing — may be protected. Most major vendors assign output rights to the user in their terms, but they typically do not warrant that the output infringes nobody else's rights. Those are two separate things. So three things are worth doing: keep a record of how a piece was made (prompts, iterations, human edits — that record is your evidence in a dispute); have core brand assets such as the logo, key visual and mascot led or substantially reworked by a person; and run a pre-publication check on brands, trademarks, likenesses and advertising claims. This is general information rather than legal advice — for your own situation, ask a qualified lawyer.
07Our products look different from batch to batch. Can this work?+
Not tier three — there's no stable feature for a model to learn, and forcing it produces something unreliable. We'll tell you that directly. Tiers one and two still hold: as long as this batch has been photographed, scenes and backgrounds can still be produced at volume.
Next