Case study · Live · Our own business

We didn't bookmark ten websites.
We wrote the desk.Images, video, audio and 3D in one interface we wrote ourselves

One service for images, another for cutouts, another for video, another for voice — output scattered across their back offices, prompts with nowhere to accumulate, and no single gate on spend. So this layer is ours: twelve workbenches, eleven job types on one queue, five domains each with its own recipe system. It isn't delivered to clients. It's what we use daily, and the clearest example of what “we write the platform layer ourselves” actually means.

Workbenches

12

Job types

11 · one queue

Built-in recipes

213 · 5 domains

Spend gate

Daily ceiling · refuses, not warns

How it started

Not a wheel worth reinventing — the pieces had scattered

None of these four is “the tool wasn't good enough”. As capabilities multiplied, the problem stopped being how to produce an image and became where everything went.

One site per capability

One service for images, another for cutouts, another for video, another for voice. Each one's output stayed in its own back office, so finding last month's image meant logging into four places and scrolling.

Output scattered across other people's accounts is not an asset library

Prompts rewritten every time

How was the last batch produced? Nobody remembers, so it gets written again from impression. The batches stop matching, and matching gets harder over time — because what has to be matched was never written down.

Style held together by memory isn't held together

No gate on spend

Video costs one to two orders of magnitude more per call than an image. Click through on “this probably isn't expensive” and the bill arrives at the end of the month. Most tools offer a usage alert, not a refusal.

An alert isn't a gate — by the time you see it, it's spent

General tools don't separate domains

The vocabulary that works for product photography is wrong for game art: one wants real light and material, the other shouldn't contain photographic terms at all. If the tool doesn't separate them, a person has to remember to, every time.

Two different things forced through one input box

What we actually wrote is only the middle layer — the interface, the queue, the recipes, the asset library and the ledger. The model capability is still called and still swappable, which is the same thing we tell clients when we say a self-written platform is not a black box.

What it is

Four figures, all countable from source and library

Which is why they're here. Throughput and hours saved have no checkable source, so there are none.

19 pages, 12 of them workbenches

Generate, model, cut out, 3D render, video workshop, video generate, music, voice, expression sheets — plus generate / video / music again for the game line. The rest are the overview, asset library, recipe library, jobs, usage and account settings.

11 job types, one queue

Images, cutouts, 3D renders, model builds, video generation, editing, music, voice-over, transcription, compositing, audio post. Which page you start from is your choice; queueing, retry and logging are the same for all of them.

5 domains, each with its own recipe system

Product, game art, music, game music, video — prompt vocabulary, defaults and aspect ratios are not shared. Things that should be separate are separate at the data layer.

213 built-in recipes

Camera, composition, light, material, style, figures, motion — organised by domain. You pick a recipe rather than rewriting a prompt.

The reasoning behind it

Seven decisions that haven't changed since

Tooling changes and models turn over every few months, but none of these seven has moved. Whether a system gets used daily rests on judgements like these, not on picking the right model.
01

The product line and the game line are two systems, not one toggle

Separate workbenches with separate prompt vocabulary, recipes, defaults and aspect ratios. Product imagery wants real light and material; game art shouldn't contain photographic terms at all. Merge them into a “style dropdown” and both sides end up compromised.

02

A recipe isn't a template — it's a record of how to ask

Every time a way of asking produces a good result, it becomes a reusable recipe rather than something in one person's head. That's why the second batch matches the first: what they match against is a stored recipe, not a memory. Recipes can be edited and added to, and one edit applies from then on.

03

First principle for product imagery: it has to be that product

Whenever there's a reference photo, the path that treats it as the source of truth is the only one used — otherwise the product itself gets repainted, and what comes out is something that resembles what you sell rather than the thing you sell. We give up freedom for that, deliberately. The cost is real: that path won't add new elements to a scene, which is why “there should be a person in the shot” had to become its own class of recipe.

04

Reference images always offer both routes: upload, or pick from the library

All seven places that take a reference image use the same picker. That makes the asset library an input, not just an archive — last week's output can feed this week's job without a download-and-re-upload round trip.

05

Where money is spent, a hard gate — not a reminder

A daily ceiling, with every call priced before it runs: if the estimate would exceed it, the call is refused, not flagged for you to decide about. Actual spend is ledgered per call and queryable by section and by day. Video costs one to two orders of magnitude more than images, so this had to work before video was connected at all — in the other order, the tuition is paid in cash.

06

Where money doesn't have to be spent, don't spend it

The whole 3D side uses no AI at all: resizing a built model, re-laying-out an exploded view — those are computed locally, so re-running costs nothing. Knowing which step is “ask a model” and which is “just calculate it” saves more than any amount of parameter tuning.

07

Never let the model write the text on an image

Let a model set type and some fraction of the time it will be blurred, misspelled or a letter long — and fixing one word means re-running the whole image. So text is composited afterwards: exact position, correct glyphs, and ten language versions of the same image in one pass. This isn't a workaround for a limitation; it's how the job should have been done anyway.

Five of the seven trade freedom for stability: only the reference-constrained path, never let the model set type, don't chase models, no pixel-level control, compute what shouldn't be asked. Each one alone looks like narrowing the tool. Together they are why a batch matches and why the month's spend is not a surprise.

Ruled out

Four things crossed off at the start

Written down because they get treated as “we'll add that later” when they are actually design premises.

No multi-tenancy It is built for one person. Concurrency is minimal, and building it as a tenanted system would buy nothing but complexity and harder debugging.

No local generative models This machine has no GPU. Generation is an API call; the machine handles only what a CPU can do — model building, exploded views, compositing, post-processing.

No pixel-level control Pixel-exact work belongs in design software. This desk produces batches reliably; the last mile stays with a person.

No chasing new models Swapping a model is easy. What's hard is whether the batch still matches afterwards. Stable is worth more than new.

Three pages, three jobs

Everything here that mentions imagery, separated

Set out so you know which one to open, rather than reading the same thing three times.
WhereWhat it coversIn concrete terms
Visual asset pipeline (service)What imagery can be producedThe imagery capability delivered to a client, in three tiers: direct generation / constrained by the real photo / a brand-specific model.
Marketing content pipeline (case)How one piece travelsProduct photography in; scene renders, short video, copy and hashtags out, packaged to whoever publishes, with a log and platform data flowing back.
This pageWhere the work happensThe desk behind both of those. It isn't delivered to clients — it's the tool we use daily, and the clearest example of “we write the platform layer ourselves”.

Questions

What you're likely to ask

Answered directly.
Isn't building your own just reinventing the wheel?+

If you use one or two capabilities, off-the-shelf is cheaper, and that is how we started. The turn came as the capabilities multiplied: images, cutouts, model building, 3D, video, voice and music sat with seven or eight different services, output scattered across their back offices, prompts with nowhere to accumulate, and no single gate on spend. What needed building at that point wasn't another image tool — it was the layer that ties them together, and no single vendor will ever build that, because it crosses all of their boundaries. What we actually wrote is narrow: the interface, the queue, the recipes, the asset library and the ledger. The model capability is still called, and still swappable.

Could you deliver this to us?+

This page describes the tool we use daily, not a packaged product. The approach is deliverable though — the structure, the recipe system, the spend gate, the library-as-input design can all be rebuilt around your categories, which falls under the visual asset pipeline service. Whether it's worth that depth depends on volume: a few dozen images a month doesn't justify a desk. Producing every week, across languages and formats, and needing them to match — that does.

Why make a point of never letting the model set type?+

Because it's where e-commerce imagery most reliably fails: blurred type, a misspelling, one letter too many — and fixing one word means re-running the whole image, with some fraction wrong every time. Compositing text afterwards makes position exact and glyphs correct, and produces ten language versions of the same image in one pass, with no re-run to change a word. Judgements like that — which step should ask a model and which should simply be computed — decide whether a system gets used daily. They matter more than which model you picked.

How does the spend gate actually stop anything?+

Every call is priced before it runs; if the estimate would exceed that day's ceiling, the call is refused rather than flagged for you to decide about. Actual spend is ledgered per call, queryable by section and by day. One ordering rule is firm: this had to work before video was connected at all, because video costs one to two orders of magnitude more per call than an image. Done in the other order, the tuition is paid in cash.

Why are there no throughput figures on this page?+

The figures here — workbenches, job types, domains, recipe count — can be counted from the source and the library, so they're publishable. Throughput and hours saved have no checkable source, so there are none. Invent one and every verifiable number elsewhere on this site loses value with it.

Next

First, check whether your volume justifies it

A few dozen images a month is cheaper off the shelf, and we'll say so. Producing every week, across languages and formats, needing them to match — that is when a desk of your own starts to pay. The diagnostic fee is credited in full against the work that follows.