Service · Visual asset pipeline

Give it your product photos,
get back a full set you can shipImagery, video, 3D and voice — all from one visual standard

Most companies have tried AI imagery and given up, and the reason usually isn't that the pictures looked bad. It's that a batch never came out as a set, and the product itself came out wrong. Both are solvable — not by finding a better model, but by picking the right tier and then fixing the process around it. Once that process holds, video, 3D and voice-over are extensions of the same standard rather than four suppliers each doing it their own way.

In

Photography you already have

Out

Imagery · video · 3D · audio

Runs on

Your own servers

Delivery

In blocks · scoped at the diagnostic

The pipeline in one line

Photography in, a shippable set out

What goes in

The product photography you already have. No reshoot, and no need to organise it into a particular format first.

What comes out

Scene imagery, sized hero and detail images, social crops and short-form video source material — consistent across the batch, with filenames and alt text generated, ready to write back to your catalogue.

The part in between

Runs on your own servers. Material, configuration and models never leave your boundary, and nobody is billing you per image.

What actually goes wrong

It isn't ugly, it's uncontrollable

These three are what stops companies in practice — and the reason the next section asks you to choose.

Different every time

The same prompt ten times gives you ten different looks. Fine for one image, fatal for a product range.

You can't specify it

“Product on the left, space above” — a text description can't control layout precisely. Every revision is another roll of the dice.

The product comes out wrong

A general model has never seen your product, so the details are invented. On a commerce detail page that's disqualifying.

And one more practical one: outsourced product imagery costs tens of currency units per image, takes days, and still doesn't match across a set. Launch often enough and artwork becomes the bottleneck for everything downstream.

Not only imagery

Four kinds of output, one visual standard

This line started as product imagery and has since been extended three times. What matters isn't that it can produce everything — it's that all four come from one standard: the product in the video is the product in the hero shot, and the voice-over says what the detail page says. Commission four suppliers and that alignment is exactly what you lose.
01

Imagery

Scene shots, cut-out heroes, detail images, social formats. Every size exported in one pass with its own safe area, margins and subject proportion — not one crop resized; filenames and alt text generated with them, ready to write back to your catalogue.

A batch arrives as a set

02

Video

Product motion, scene shorts, vertical social cuts. Built from the same approved images and the same visual standard, so the video and the stills don't contradict each other.

One step after the imagery

03

3D

A usable three-dimensional model from product photography or a description, then renders from any angle, structural breakdowns and exploded views. Resize it, re-lay-out the breakdown — that part is computed on your own machine, so another version costs nothing extra.

Another version costs nothing

04

Audio

Voice-over, subtitles, background music. Every language version comes from the same script, so the wording doesn't drift in translation.

One script, every language

You don't need all four. Most engagements start with imagery and extend once that runs — the foundation is shared, so each class after the first is noticeably lighter.

Three tiers

Not one method — a choice between three

Which tier you need depends on how exactly your product has to look. This is the section worth reading twice: pick wrong and either the images are unusable, or you paid for accuracy you never needed.
Tier 01Fastest · cheapest

Straight generation

A general model generates the image outright. Suited to mood shots, backgrounds and social filler — anything where the product doesn't have to be exactly right. High volume, fast turnaround.

The limit: the product's actual form is guessed. Not acceptable for a hero or detail image.

Tier 02Where most commerce imagery lands

Constrained generation

Your real photograph is the reference the output is bound to: the product itself isn't repainted, only the background, lighting and setting are rebuilt. Framing and silhouette hold, a batch comes out as a set, and it can go straight onto a product page.

The prerequisite: the product needs one usable photograph to work from.

Tier 03Yours to keep

Brand-specific model

A model trained on your own product photography, so it knows your products specifically. After that you no longer need a reference shot every time to get the form right — including angles and settings photography can't practically reach. Worth it with many SKUs, frequent launches, or heavy demand for scenes you can't shoot.

The model is your asset. It lives on your servers and you can take it with you.

The three mix. Hero images at tier two where accuracy is non-negotiable, scenes and social at tier three for volume, mood material at tier one. We give a recommendation per category during the diagnostic — including “tier two is enough, don't pay for tier three”where that's the honest answer.

How a batch gets through

Four stages, one of them yours

What follows is who is doing what.
01You hand over source materialyouProduct photography, plus your visual standard — the tone, how much negative space, which export sizes. If no standard exists, we set one once and it's fixed after that.
02The batch is producedmachineProcessed as a batch, several candidates per image. It checks its own output first and re-runs what fails, which is time that doesn't come out of your week.
03Your people chooseyour teamSomeone picks the final version. This step is deliberately not automated — it's a design decision. Rejected frames go back for a re-run.
04Delivery is automaticmachineEvery size exported at once, filenames and alt text generated, written back to your image library. Who approved which batch and when is all logged.

Your own team runs the next launch without us. For scale: selection takes us roughly 20 seconds an image, about eight minutes for a batch of 24 — but that's our own pace on our own catalogue, not a commitment about yours. Your batch sizes, categories and review standards differ, and we estimate against them during the diagnostic.

Video and audio

The same material, extended into video and voice

Once the imagery exists, video source material is one more step, not another project: product motion, scene clips, vertical cuts for social. It uses the same approved frames and the same visual standard, so the video and the stills belong to each other instead of arguing. Voice-over, subtitles and background music follow— every language version comes from the same script, so the wording doesn't drift in translation.

What's delivered is repeatable production capacity, not a folder of finished clips. Next month's launch runs the same pipeline with your own people. That's the difference from commissioning a batch of videos: that budget is spent and gone, this one leaves something behind.

Not included in this part

  • ✕Physical photography
  • ✕Voice talent and on-camera models
  • ✕Scripting and creative direction

3D

Some products can't be explained by a flat image

Furniture, equipment, fittings — what a customer actually wants to know is how it assembles, what's inside, and how much room it takes, and a photograph struggles with all three. Once a model exists, the angles, the structural breakdown and the exploded view all come from it, so you never end up with five screws in the diagram and six in the box.

Renders from any angle

One model, any viewpoint — including the ones no camera could be rigged for. Changing the angle doesn't mean scheduling another shoot.

Breakdowns and exploded views

How the parts go together and in what order, in one image. Useful for assembly instructions, support tickets and dealer training alike.

Sizes and configurations

Different sizes and specs in the same range: change the parameters and re-output. That part is computed on your own machine, so another version costs nothing extra.

Two boundaries on this one

One: whether it comes apart depends on how the model was made. A model converted straight from a photograph is usually a single shell and won't separate into parts; an exploded view needs a model built structurally. That gets settled at the diagnostic, not discovered before delivery.

Two: this is not engineering drawing. It shows what a thing looks like and how it goes together, not tolerance-accurate parts you could tool from. For that precision, you want a structural engineer.

Brand-specific model

What training one asks of you

What you supply, what you get, and when not to bother. How it's trained is a proposal-stage conversation.
Source material20–40 photographs per product, covering the main angles and different lighting. Material quality sets the ceiling, and there's no way to skip this.
TimeHow complete the material is decides the pace — selecting and labelling images is the bulk of it, not the training itself. Scheduling is set during the proposal against your categories and volume.
ValidationWe test the finished model against images it never trained on and check the form and key structural parts. If it doesn't pass, the material changes and it trains again. We don't ship a model that's roughly right.
OwnershipTraining material, configuration and the model file all sit on your servers. It's your asset and you can keep training new versions yourself.
When it doesn't applyProducts whose appearance changes batch to batch — there's no stable feature to learn. We'll say so rather than force it.

Three that don't bend

Each one narrows the tool. Together they make the batch usable.

Models turn over every few months; these three haven't moved. Alone each looks like giving up freedom — together they are why a batch matches and why the month's spend isn't a surprise.
01

It has to be that product

If a customer will use the image to judge what the thing actually looks like, it is generated from your own photograph as the source of truth and the product itself is never repainted. We give up freedom for that deliberately — something that merely resembles what you sell comes back to you as a return.

02

The model never sets type on an image

Let it, and some fraction of the time the text is blurred, misspelled or a letter long — and fixing one word means re-running the whole image. So text is composited afterwards: exact position, correct glyphs, ten language versions of one image in a single pass, and changing a word re-runs nothing.

03

Spend has a gate, not a warning

Every call is priced before it runs, and one that would exceed the day's ceiling is refused rather than flagged for you to decide about; actual spend is ledgered per call and queryable by section and by day. Video costs one to two orders of magnitude more than an image, so this has to work first.

These came out of the desk we built for ourselves; the full set is on our own visual production desk →

What you end up with

A pipeline that runs itself, not a folder of images

This is the difference from hiring a retoucher: next launch, you run it.

A repeatable pipeline

The same configuration next month produces a set that matches this month's. Nobody has to remember how it was tuned.

Your brand-specific model

Delivered when you go to tier three: trained on your product photography, stored on your servers, dependent on no platform.

An export standard

Sizes, spacing and naming rules for hero, detail and social, written down and handed over so a new hire can follow it.

A run log

Who approved each batch, when, and how many frames came out — traceable, and resumable from where it stopped.

Not included

  • ✕Physical photography: materials, real usage settings, models
  • ✕Models, locations and location shoots
  • ✕Voice talent, on-camera models, scripting and creative direction
  • ✕Holding your ad accounts or making budget decisions
  • ✕Content compliance and legal review

When it isn't worth doing

  • ✕A few launches a year and a few dozen images total — outsourcing is cheaper
  • ✕No brand visual standard yet: a machine executes a standard, it can't decide one for you
  • ✕Products that look different every batch — there's no stable feature for a model to learn
  • ✕You want a handful of striking campaign images — that's creative work, not pipeline work

What this looks like wired into a storefront: E-commerce automation & support agent →

Questions

What people ask about this one

The first is the selection question, which is the one that matters most here — so it's first.
01Which tier do we need?+

One question decides it: will a customer use this image to judge what the thing actually looks like? If yes — hero images, detail pages — start at tier two, where the product's form has to be correct. If no — mood shots, backgrounds, social filler — tier one is enough and anything more is wasted money. Tier three earns its cost when you have many SKUs, launch often, or need volumes of angles and settings photography can't reach. During the diagnostic we give a per-category recommendation, including “tier two is all you need” where that's the answer.

02Can the output go live as-is?+

Yes, but it passes one of your people first. Selection is the step we deliberately don't automate — a hero image moves your return rate, and nothing in that category should publish without a human having looked at it. The machine prepares the candidates and every size; which version ships is your call.

03Do we need a brand visual standard first?+

Ideally yes. If there isn't one we'll set it once and write it down, and it stays fixed after that so a new hire can follow it. But to be clear about the boundary: a machine can execute a standard, it can't decide what your brand should look like. Without that decision, the output can only be attractive — it can't be recognisably yours.

04Is the custom model a recurring cost?+

No. Training material, configuration and the model file are all on your own servers and belong to you; you can keep training new versions or hand it to someone else. We don't bill per image or per call — how much you generate doesn't change what you pay us. The other half should be just as clear: generating images does cost money. That cost isn't in the delivery fee and doesn't pass through us — you top up your own account, the compute runs on your own servers, and we take no cut and add no margin; spend is visible per module, per person, per month. What one image costs depends on the model, the output size and how many passes it takes: on our own line it works out at around $10 a shot, against $30 outsourced (source: image production log). Your categories and sizes differ, so the per-image figure and a realistic monthly total are worked out against your actual volume during the diagnostic.

05Does AI-generated imagery need to be labelled?+

In most markets now yes, and some of it is already in force. China's labelling measures for AI-generated content took effect on 1 September 2025, requiring both visible labelling (on-image marks, text notices) and embedded labelling such as file metadata. The EU AI Act's transparency obligations apply on a timetable from August 2026. Several ad and social platforms have their own disclosure rules, and non-compliance can get an ad rejected or throttled. For this line it comes down to one thing: the label should live inside the pipeline rather than depend on somebody remembering to add it — marks, notices and metadata can all be written in at export, and that part is in scope. What is not in scope: deciding how your market and your channels specifically require you to label. That is a compliance question, the rules differ by region and keep changing, so work from current local requirements and your own counsel.

06Who owns the imagery it produces?+

Positions differ by jurisdiction, so only the broad shape is worth stating: output generated almost entirely by AI, with little human authorship, is generally hard to claim rights over; the human contribution — prompting, selecting, arranging, editing — may be protected. Most major vendors assign output rights to the user in their terms, but they typically do not warrant that the output infringes nobody else's rights. Those are two separate things. So three things are worth doing: keep a record of how a piece was made (prompts, iterations, human edits — that record is your evidence in a dispute); have core brand assets such as the logo, key visual and mascot led or substantially reworked by a person; and run a pre-publication check on brands, trademarks, likenesses and advertising claims. This is general information rather than legal advice — for your own situation, ask a qualified lawyer.

07Our products look different from batch to batch. Can this work?+

Not tier three — there's no stable feature for a model to learn, and forcing it produces something unreliable. We'll tell you that directly. Tiers one and two still hold: as long as this batch has been photographed, scenes and backgrounds can still be produced at volume.

Next

Run a batch of your own images first

During the diagnostic we run a batch of your actual product photography and tell you which tier to stop at, what belongs in a pipeline and what still has to be shot — including the parts not worth doing. The findings are yours either way, and the fee is credited in full.