ArtCraft and the Case for Treating AI as a System, Not a Black Box

October 9, 2026

Over the past week, a developer named Brandon Thomas released ArtCraft: seven open-source, clean-room Rust reimplementations of Adobe's flagship apps — PhotoCraft for Photoshop, FilmCraft for Premiere Pro, VectorCraft for Illustrator, and four more — built largely with Claude Opus. The coverage has mostly focused on the headline: "AI cloned the Adobe suite." That framing is a little misleading, and also the least interesting part of the story.

The more useful story is about what Thomas did around the AI, not what the AI did on its own.

The black-box failure mode

The default way most people use a coding model looks like this: ask it to build something, read the diff, ship it if it looks right, and take its word for how done the feature is. The model becomes an oracle. You trust its account of its own work because checking that account is expensive, and the model is persuasive by default.

ArtCraft's documentation is refreshingly allergic to this. A few examples:

It stopped trusting self-reported progress. The project's own roadmap admits that many of its "percent complete" numbers were written by the same agents doing the work, and flags them explicitly as "not checked independently." Where it could measure instead of ask, it did: PhotoCraft's menu-wiring is 100% complete, but its actual pixel-match rate against real Photoshop files on a test corpus is around 68%, and a separate estimate of correctness-in-depth lands around 60–70%. Three different numbers, from three different kinds of ground truth, deliberately kept apart instead of collapsed into one reassuring headline figure.

It gave agents an interface to be watched through, not just a codebase to write into. Every user-facing action in these apps is implemented as a single registered command — with an id, a shortcut, validation, and a test — and the UI, CLI, and an MCP server all dispatch through that same registry. That's not an architecture flex; it's what makes an agent's behavior observable and replayable instead of a trust exercise. You can drive the app the same way whether a human is clicking or an agent is issuing commands, and you can watch exactly what happened.

It treated "many agents working at once" as a coordination problem, not a productivity multiplier to lean on blindly. Parallel agents were assigned to separate subsystems with their own build directories and explicit rules about which files they could touch — and a human still did the integration, resolved the conflicts, and decided what shipped. The scaling wasn't "more agents, less oversight." It was "more agents, with oversight moved to the seam where their work meets."

It published the gap instead of smoothing it over. The repos carry open bug trackers, a public banner noting the apps aren't ready for daily professional use, and roadmaps that separate "features exist" from "features are correct." That candor is itself a design choice — it only works because the project built the measurements needed to be candid about.

Why this matters more than the Rust rewrite

None of this made ArtCraft's agents smarter. What it did was make their output falsifiable. A black-box approach to AI coding treats the model's claim and the actual state of the software as the same thing. ArtCraft's approach keeps them separate on purpose, so disagreements between them are informative instead of invisible.

That's the actual shift happening in how serious teams use coding agents right now: less effort going into better prompts, more effort going into building the scaffolding — interfaces, oracles, fuzzers, test corpora, explicit coordination rules — that let you verify and constrain what an agent did, rather than just read its summary and move on. The AI isn't the system. It's a very fast, occasionally wrong component inside a system someone still has to design.

The Adobe clones are a fun demo. The willingness to measure them honestly is the part worth stealing.