Why general-purpose image models can't produce manufacturable jewelry
mulapis Research
A general-purpose image model optimizes for visual plausibility: the scene looks real, the light looks real, the material looks real. Jewelry production instead demands that geometry and craft hold up. The two objectives are not the same — and that divergence is exactly the distance between a beautiful image and a producible design.
An image has no geometry
An image model outputs pixels, not structure. A render contains no wall thickness, no dimensions, no seat depths — the information production requires simply does not exist at the pixel level. Manufacturing from an image means asking the factory to redesign the piece by eye.
Visually plausible is not craft-sound
From millions of photos, an image model learns what jewelry looks like — but never the physics of casting. It draws floating settings, prongs thinner than the casting minimum, openwork with no tool access — each plausible on screen, none workable on the bench. These are not random flaws; the model's objective simply does not include manufacturability.
Production is a constraint problem, not a style problem
Handing jewelry design to AI is really a constrained-generation problem: metal, stones, setting, ring size — every field of the spec is a condition the generation must satisfy. mulapis approaches this with a fine-tuned orchestration model dispatching the work and a specification review validating each item, with failing designs redone automatically — the constraints live in the process, rather than hoping a single model learns them all.
Revisions must be controllable
Real orders run on local revisions: resize the band, keep the stone; swap the side stones, keep the shank. An image model regenerates the whole picture each time and cannot guarantee that one change leaves the rest untouched. Controllable revision requires a design carried by spec and geometry, not by pixels.