Recompile, don’t refactor. How I build production software with AI: not the app, but the system that compiles it
- Don’t build the app. Build the system that compiles it. The real prompt is the coding harness: code rules, a design system, a roadmap with specs, and a data model. The only prompt I wrote was “Build the whole app from the roadmap.”
- The app is disposable. Feedback and scope changes go into the harness, not the code, and the next build already has them baked in. The app was regenerated from scratch ten times, and I didn’t touch the code once.
- Big feedback costs as little as small feedback. The cost of a change isn’t proportional to the code that already exists, but to the lines added to the harness. When a phase-two module broke the data model, the fix took a day, and then a new build ran.
- The client tests from week one. The first version of the app is always just a frontend, indistinguishable from production. Four iterations with the client validated the product, the scope, and the data model before the backend existed. The code from this phase was a by-product.
- A component library instead of a tokens doc. The agent never followed a tokens doc. Vibe coding is only there to find the look. The agent extracted a library from that look and maintains it on its own through the 80 percent rule, derivation recipes, and a design audit.
- Ten hours, one prompt, zero questions. The orchestrator runs the agents epic by epic through a verification gate and two audits, and never trusts that “tests passed.” Result: 64 pages, 947 tests, and fifty clicked-through journeys without a blocker.
- Go one level up. Six weeks instead of twelve, one person instead of a team. Only a bug that’s missing a rule deserves a rebuild. The reward is no longer a piece of code that works, but a system that generates it right. And this harness is disposable too.
AI wrote this summary from the full article. The author checked it.
I don’t write code anymore. I change the product definition at the top, and coding agents recompile it into clean application code in a few hours, again and again, with no technical debt. This is how I build production apps for clients, from scratch, built entirely with AI.
Four days before the deadline, all I had was a system that compiles. No app. Before agentic engineering, by this point you’d already be a few weeks into testing and debugging the app. Here the design was done, but the app’s code didn’t exist. I kicked off a ten-hour production build with AI, and by evening we had a finished production app for the client. It was quite an experience: living on the bleeding edge, where you trust the system you built, not the code you see.
For over seventeen years, I’ve been architecting, designing, and coding apps with business owners. After four years building apps in a team as a product manager, I got back into a delivery role: I could once again own the end-to-end process myself, think hard about AI, and reinvent how an app gets made. This article is about how I design and build greenfield production apps for clients today, built entirely with AI: one client project, from nine workshops to a ten-hour production run.
#Three things that changed
I didn’t go the vibe coding route, where you build the app one prompt at a time and every change is another patch. Without an architecture decided at the start, at some point it turns into chaos and technical debt that stops your project in its tracks. Instead, I built a system that reinvents three things software development has stood on for seventeen years. In this process, they’re game changers.
-
The app is disposable. I don’t write any code. I change the product definition at the top, and in a few hours the agents recompile it into clean code from scratch, with no technical debt and every change applied. On this project, the app was built from scratch ten times, and I didn’t touch the code once. A scope change that broke the data model, and would have cost a week of refactoring in code, cost a day and one build here.
-
The client tests from week one, not at the end. The biggest risk of traditional development: the product is almost done, and only during testing does the client find out that something is missing or badly designed. Here, from week one, the client had an app in hand that was indistinguishable from production, in almost full scope. They gave feedback and saw a new version the next day. Four times. By the time the real production app was built, the product, the scope, and the data model had already been validated.
-
One person, one prompt, ten hours. One autonomous agentic run on my computer built the whole production app, from one prompt, without a single question. Executing the whole project took six weeks. Before the era of agents, the same scope would have taken twelve weeks and needed a product manager, a designer, and two to three engineers.
About the app: the client, Marco CAR, an automotive company in Slovakia with fewer than 50 employees, commissioned an app from me to run its internal operations, plus a customer-facing module where business customers manage their own fleets. Repair cases, vehicles, appointments, service history, notifications. Four user roles, each with its own interface, eight modules, two phases.
The rest of this article is how I built it all.
#Before the first build
##Nine workshops
Before the build itself came nine workshops with the client, ~90 minutes each. The goal was to understand the requirements, explore the company’s business processes, and clearly name the problems its custom software should solve. We worked on two tracks: designing new internal operations and the software to run them, and a customer-facing module of the same software, because the company cares about its customers’ long-term satisfaction. After the workshops, I knew what to build.
From the very first workshop, AI helped me store the information and give it a clear structure. Throughout the project, that structure was the main source of truth for decisions about the system’s design.
##Two harnesses
This is a good place to introduce two terms that will keep coming up in this article. The project has two harnesses, and each one does something different.
The knowledge harness is my own system for organizing the information from all the workshops with the client and synthesizing it into the final outputs: the roadmap and the feature specs. It’s also the project’s operating system: meeting notes, decisions along with their reasons, a map of the people on the client’s side, all the communication. From that same state, it puts together my meeting materials, an email to the client, a proposal, and a contract, and it answers the question “why did we decide on this?” with a link to the specific note. How it works on the inside I’ll save for a separate article.
The coding harness is a set of text documents in the app’s repository that define the rules for the coding agent: how to write the frontend, how to write the backend, what the architecture is, how to test, and how to deploy. It’s the prompt the agent gets instead of a brief. The rest of this article is about it. From here on, when I write just “harness,” I mean this coding harness.
The relationship between them runs one way: the workshops go into the knowledge harness; out of it come the roadmap, the specs, and the data model; and those go into the coding harness as the brief for what to build.
##Tech stack
Let’s look at my tech stack:
- Agent – Claude Code, extended with my own skills for synthesis and management in each phase of the project: discovery, roadmap, spec, build, design/code audit.
- Frontend – React 19 with TypeScript and Tailwind v4, built with Vite, no UI library: every component is custom, and I’ll explain why when we get to the design system. In the design phase it ran on its own, without a backend, with IndexedDB as a local database.
- Backend – Laravel 13. The frontend sits on top of it through Inertia, so the app is a monolith with no API layer: a controller sends data to the page as props, and a form sends it back.
- Communication and integrations – Emails through Laravel Mail and Notifications (account activation, monthly reports, reminders), via Resend in production. SMS through BulkGate, with no package, just a thin custom channel over the HTTP client. PDF through dompdf to generate internal documents. The scheduler for month-end closing and appointment reminders. Integration with SoftApp, the service shop’s inventory and invoicing system, with service records matched by license plate.
- Infrastructure – PostgreSQL, Redis with Horizon for queues, Fortify for authentication, private object storage for invoices and statements, monitoring through Pulse and Nightwatch, tests in Pest, hosting on Laravel Cloud.
One rule in the harness is worth mentioning: every problem is solved with a first-party Laravel package, and nothing else gets installed without my approval and an entry in the decision log. That leaves the agent no room to improvise with libraries.
#By the numbers
#How do you generate an app ten times so it looks and works the same every time?
And on top of that, with the client’s feedback built in each time. Without an answer to this question, you can’t iterate like this. You don’t want to hand the client the app for the tenth time and have it look different every time.
The answer is the coding harness, our compiler. Everything that has to stay stable between builds is defined at its level. The design system: how the app looks, what interaction patterns and rules it has, and how to derive new components that aren’t in it yet. The code rules: exactly how to write the frontend, how to write the backend, what the architecture is, what the testing strategy is, how tests are written, and how deployment works. And finally, the backbone of the app: the full scope of features and a spec for each one. Everything that matters is defined once, at the harness level.
The next three chapters walk through these layers: the design system, the frontend rules, and the roadmap with its specs.
#The app’s design system
I wanted to reinvent how a brand-new product comes to life, from the first idea to production. So let’s start at the beginning: how I created a design system without opening Figma and starting to draw screens.
##Vibe coding to find the look
I already had a picture in my head: what dashboard the company needed and what visual language it should speak. I focused on functional design, not aesthetics, and on that basis I started to vibe-code a dashboard for the business owner. (The examples below are real screens and the component catalog from this project, published with the client’s consent. It’s only part of the app, not the whole product.)
This is exactly where vibe coding fits in my process: finding the look, not building the product. It was nothing complicated, and you don’t need to think about best practices for the frontend stack. A new folder, Claude Code open, and a prompt to create a basic React project with the owner’s dashboard page in it, using the colors, fonts, and design direction I gave it. I wanted an admin dashboard with a navigation sidebar on the left and typical admin content on the right: tables, tabs, search, a few stats.
I then iterated on the first version and polished it several times until I liked how it looked. The basic dashboard was maybe 15 to 20 prompts. And it was interactive: the components responded to clicks. It wasn’t a static screen.
That’s how I ended up with four dashboard variants in total, each with a different kind of look. I picked the one I liked most for the project’s purposes and derived three more pages from it: an entity detail with information, tables, and action buttons; a list of entities in a table; and a form to create a new entity. These four pages contained most of the interaction elements that repeat across the app as its basic building blocks. And I could show them to the client and ask for feedback on both the look and the interaction patterns before a single real feature existed.
##From four pages to a UI component library
It was two days of manual work exploring and building the foundation of the app’s design. The next step was the key one: I had Claude extract every UI component from those four vibe-coded pages and turn them into a full-fledged UI component library for the app. To make them visible and testable, I also created a Showcase page where you can browse them all, with descriptions: fonts, colors, whitespace, and what each component is for and how to use it.
The library has one purpose: to make the coding agent build the frontend consistently across the whole app, while I keep full control over how it looks. I also tried a simpler route, a single design.md document with tokens: colors, fonts, paddings. It didn’t work. The agent never stuck to the look, and in every run it derived a different kind of app, never a consistent one. A custom component library, finished before the first build, turned out to be the only reliable way to get the same frontend across ten runs.
A few numbers, to make it tangible. From those four vibe-coded pages, Claude extracted 46 components. Today, after the production run, the catalog has 83: 63 primitives (buttons, fields, pills, cards, tables, drawer, modal), 4 composites (app shell, sidebar, page header, form card), and 16 domain components like the case timeline or the service history table. I didn’t write a single one of those 37 new ones. The agent derived all of them during the builds.
##Why no third-party UI library
This is the right place to say why the project has no third-party UI library. I’ve never used them. I’ve always preferred to design and code my own components. For me, UI libraries are limiting both aesthetically and functionally: they bring chaos and hidden complexity into the code and push their own opinionated direction on how an interface should look. I like clean, custom solutions. And today it’s no longer a trade-off between speed and cleanliness: a coding agent codes up a component with everything we need, on the spot, with none of the baggage a library would bring along. The deliberate trade-off is accessibility, which ready-made libraries like shadcn/ui handle for you (keyboard, focus, ARIA...), but it’s nothing the agent can’t manage when the project calls for it.
##The design contract: how the system is used
Once the components were coded, I wrote a design contract into the harness. It’s not a description of the look. That already lives in the code and on the Showcase page. It’s a manual for how the system is used and where its boundaries are. It has seven parts:
- Agent contract – six steps in a fixed order that the agent goes through for every screen: pick a page template, pick components by purpose, derive only when nothing fits, never invent tokens, put every new component in the Showcase, stick to the localization.
- App shell and canvas – one canonical shell for all roles, no top bar, white canvas, sepia only in drawers, fixed sidebar dimensions, full-width content. Three page templates (dashboard, list + detail, centered form) that every new screen is derived from. A drawer only on the dashboard for a quick preview, a modal for one focused question. A single breakpoint where the shell switches to mobile mode.
- Component selection – an intent → component table: primary action, secondary action, status pill, search, form field, KPI tile, empty state, error. The agent first names what it wants and only then reaches for a component. Plus rules for tables: column order, row actions, server-side sorting.
- Derivation recipes – eight shapes (button, field, toggle, pill, tile, surface, overlay, composite), and for each one a procedure for what to derive a new variant from and how.
- Icon contract – size and stroke weight by context, colors, bi-tone active icons in the navigation.
- Anti-patterns – a list of things that made it into the code at least
once, in this project or a sister project: no raw hex, no new accent colors,
no orange buttons, no heavy shadows, no
font-bold, no third-party UI library, no made-up variants. - When this document is wrong – if the rules block a reasonable solution, that’s a signal to escalate, not permission to break them.
| Looking for… | Go to… |
|---|---|
| Color, radii, shadow, typography tokens | src/index.css (@theme block) |
| Live, runnable showcase of every primitive | src/pages/DesignSystem.tsx (route /design-system) |
| One-line catalog of every component + props | docs/COMPONENTS.md |
| Layer boundaries / import rules | docs/BOUNDARIES.md |
| Project architecture, hooks, services | docs/ARCHITECTURE.md |
| Routes & navigation | docs/ROUTES.md |
| Error UX pattern | docs/ERRORS.md |
| Intent | Primitive |
|---|---|
| Hero action on a page | PrimaryCTA (dark fill). Never orange. |
| Cancel / secondary action | SecondaryButton |
| Toolbar action with icon + label | GhostButton |
| Icon-only button (with aria-label) | IconButton |
| Status label | Pill (semantic tones) — or StatePill for case lifecycle |
| Hero search bar (Dashboard) | HeroSearch — live result list inside the card |
| In-table / compact search | TableSearch (inside TablePanel) |
| Form field | TextField (mono / prefix / suffix / error all built-in) |
| Two/three-option choice in a form | SegmentField |
| Yes/no setting | Switch (ink default; orange only for network-level settings) |
| Card list of options where one is selected | RadioGroup |
| KPI numbers above a table | StatTile (only one highlight tile per row) |
| Container with elevation | Card |
| Container without elevation | Surface (tints: surface, base, subtle) |
| Filterable list | TablePanel + DataTable + TableRow + TableCell |
| Empty list | EmptyState |
| Inline notice | AlertBanner (tones: danger, warn, info) |
| Record detail opened from a list page | Full-page detail route — BackLink + DetailPageHeader + DetailCards. Never a drawer (see §2.6) |
| Quick record preview or the activity panel — Dashboard only | Drawer (right-aligned overlay) — see §2.6 |
| Single-question edit | Modal (centered, ~480px) |
| Numeric ± edit inside a modal | Stepper + DiffStrip |
| Inner-page title block | PageHeader + BackLink (list / form pages); DetailPageHeader + BackLink (entity profile pages) |
| Section divider with accent | SectionAccent + SectionLabel, or SectionHeader for the full block |
| Activity feed grouped by day | TimelineFeed |
| ID / serial number / installation number / contract number / shortcode | wrap in <Mono> |
| Loading state | Skeleton (shape-matched) for surface fills, Spinner for inline, PrimaryCTA loading for buttons |
| Failure state | ErrorDialog (centered modal — never inline error rows; see docs/ERRORS.md) |
| Context | Size | Stroke |
|---|---|---|
| Nav rail / sidebar item | 20 | 2 |
| Inline with text / list rows | 16 | 1.75 |
| Ghost button leading icon | 14 | 1.75 |
| Hero search / page-level | 18 | 1.75 |
| Big empty/error state | 22–24 | 1.5 |
##Self-audit: the agent checks its own design
The contract alone isn’t enough. Coding agents drift: when they build a piece of frontend or a whole app, after a few screens they start working from memory instead of reading the Showcase. A “Create” button shows up in a top bar that doesn’t exist, the canvas turns sepia instead of white, a form makes up its own width. Every one of these actually happened. That’s why the harness has an audit skill.
At the end of every epic, the design audit runs as a separate subagent. It loads the design contract, the design system tokens, and the Showcase page, then goes through the screens that were built and compares them against those: shell and navigation, page templates, canvas color, tokens and anti-patterns, component selection, icons, localization, and whether every new component has a demo in the Showcase and a row in the catalog. Every finding has a file, a line, and the rule it broke. Without a citation, a finding doesn’t count. The audit doesn’t fix anything itself, it only reports, at three levels: blocking, should-fix, nit. The fix comes in the next step, and the epic isn’t done until it has zero blocking findings.
##Keeping the design system up to date
And the answer to who maintains the design system and the Showcase: the agent, because three rules force it to, and the audit checks them.
- The 80 percent rule – For every piece of UI, it has to find the closest existing component, and if it’s at least an 80 percent fit, it uses it and tweaks it through props. It may not fork it.
- Derivation recipes – When nothing fits, it follows the recipe for that shape: which sibling it inherits the API from, which tokens it may use, where to save the file. A new overlay can only be a drawer or a modal. There is no third kind.
- “Show what you built” – Every new component gets, in the same commit, a
demo in the Showcase and a row in the
COMPONENTS.mdcatalog. The audit flags every one that’s missing them, so the catalog never falls out of sync with the code.
The anti-patterns section of the contract came straight out of these audits: every line in it is a bug from the calibration builds that the agent actually produced and that never happened again.
The result: when the agent built a layout or UI, it always knew how. Not once did it create UI at its own discretion. It always followed rules defined up front.
We have the design system. Now let’s look at how to define the harness so it generates a consistent frontend.
#The app’s frontend
##An app without a backend
With the design system done, we can move on to defining the frontend. Let’s leave the backend aside for now. We want to get the app into our hands as soon as possible, try it out, and iterate on what we see. At this point, a backend only adds complexity and time: time we’d have to spend defining it and the agent would spend building it.
But how do you get a working app into your hands without a backend? The trick is to simulate the backend on the frontend, which coding agents handle with ease. That’s why the first version of every app I build never has a backend: it’s frontend only, and the browser’s local storage serves as its database. On the web, that’s IndexedDB. Every feature that would otherwise live on the backend lives on the frontend for now. We don’t worry that it’ll have to be moved one day: once we sign off on the final shape of the app after X iterations, we add backend rules to the harness, throw the old app away, and start a new build with a backend – as easy as swapping a printer cartridge.
##The harness, file by file
What does all this look like in the coding harness? Let’s go through it file by file. In the early phase of the project, when we were building only the frontend, the docs/ folder held these documents with predefined rules:
ARCHITECTURE.md– Describes what parts the app is made of and how they talk to each other: what a screen is, what logic is, and where data is stored. It also works as a living list of everything that gets built in the app over time, so both the agent and a human always see the current state.BOUNDARIES.md– Sets the boundaries between the parts of the app – what may depend on what and what may not, so the code doesn’t get tangled. It also lists technologies and approaches we deliberately don’t use in the project, so the agent doesn’t even start adding them.COMPONENTS.md– A catalog of all the interface building blocks (buttons, cards, tables, form fields) with a description of what each one is for. The agent is meant to assemble new screens from these pieces like building blocks, instead of inventing something new every time.DECISIONS.md– A log of important decisions: what we decided and why. Thanks to it, nobody later questions or accidentally reverses a choice that was made for a reason.DESIGN_SYSTEM.md– A binding manual for building the app’s look: how to approach designing a new screen, which visual rules are fixed, and what’s forbidden. It makes sure every screen looks as if the same designer drew it.ROUTES.md– A map of all the app’s screens and the navigation between them, including which user role sees what. The agent uses it to know where a new screen fits and who gets access to it.ERRORS.md– A single rule for how the app communicates errors to the user: always the same way, clearly, and without surprises. That way the user gets the same, predictable experience with every problem.
Take the list as inspiration, not a prescription. Every project needs different rules and different things to enforce on the agent.
These files are the alpha and omega of the rules the agent has to follow when it builds the app autonomously. I never put them in the agent’s main rules file, like CLAUDE.md, so as not to overload it. A better practice is to keep just an index of the files in CLAUDE.md or AGENTS.md: the agent reads the one that relates to the problem it’s solving right now, whether that’s architecture, new UI, or something else.
##A simulated backend and the seam under the hooks
Now back to simulating the backend. What matters is what it looks like inside. It’s not a pile of mock data in components. It’s a full-fledged database in the browser, written exactly the way it would look on the backend. The IndexedDB tables are a one-to-one copy of the data model from the harness: same entities, same names, every record has a UUID as its primary key, and relationships go through foreign keys. Even what can’t really happen in a browser, like sending an email or an SMS, gets simulated. Each send is logged as a record in the notifications table, and the link that would otherwise arrive by email shows up right on the screen. The case timeline, for example, works exactly the way it will for real, and the future backend has a seam ready to plug into (in reality it all gets rewritten, but who knows... it’s a test of good architecture design).
To keep the simulation from leaking into the UI, the harness keeps the layers separate, and BOUNDARIES.md enforces it as an import rule: types, then the database, services with the business logic above it, hooks above them, and only on top of that, components and pages. A component or page may never import the database or a service. All data flows through hooks. This was written into the harness decisions from day one, for a single reason: when the backend arrives, the layer under the hooks changes, not the screens.
#Defining what gets built
##Roadmap and epics
We have the design system and the frontend rules. What’s left is the most essential part: what we’re building. That’s what the roadmap and its feature specs are for. The roadmap looks at the app from the top and defines the modules (and therefore epics) that need to be built to get a working app. Each module breaks down into several features, and each one has a description of how it works.
The scope came out of the nine discovery workshops from the intro. The roadmap built from it has eight epics with 64 feature specs. Seven epics belonged to phase one. The eighth, fleets, came in phase two, but its scope was researched in those same workshops. No others were held.
##Specs: short and about the product
Each feature has its own .md document that describes how it should work. It’s not a classic specification like the ones you know from the early spec-driven frameworks. Mine has three parts: a Goal of three to five lines, two to three acceptance criteria, and a Layout & Design section on how the UI should be put together. For a purely backend feature, that third part is left out. People usually put more into feature specs: the data entity, its inputs and outputs, a technical specification, how to program the whole thing. I don’t do that. All the rules for how to implement something are written down once, in general terms, in the harness. There’s no need to repeat them for every feature. Today’s language models are smart enough to implement a feature on their own. Your job is to give the agent good rules at the coding harness level. Leave the rest to it.
That keeps the specs as simple as possible, so they’re easy to review, with no unnecessary detail. But one thing must never be missing: every feature and every screen states which user role it applies to. The agent won’t figure that out on its own, and in an app with four roles, “who can see and do what” is the most common source of bugs. That one sentence later grows into both the authorization tests and the user journeys.
| ID | Feature | Role | Description | Status |
|---|---|---|---|---|
| F021 | All cases list | Admin | Every case in one place. Filter by status and type. | Built |
| F022 | Case detail | Admin | Full view of one case with its timeline; everything editable while open, including swapping the case type mid-flight. | Built |
| F023 | Mark case done | Admin | Closing dialog that captures the total repair price and a short repair summary, then flips the case to done. | Built |
| Route | Screen | Portal |
|---|---|---|
| /admin/cases | All cases list | Admin |
| /admin/cases/:id | Case detail; the mark-done dialog opens from here | Admin |
##The data model as a finished input
These files come with one more document, which belongs to the product, not to the code rules: the data model. It’s a single file, entities.md, that describes all the app’s entities along with their types. It’s over a thousand lines long, and it took shape during the workshops in the knowledge harness, long before the first build. Just like the roadmap and the specs, it went into the coding harness as a finished input, not something the agent would make up during the build.
Every entity in it has the same structure. For example, a case:
- Purpose – one sentence on what the entity is for in the system.
- Attributes – a table of fields, each with its type, whether it’s required, which workshop or decision it comes from, and a note on what it means.
- Lifecycle – the states and the transitions between them: pending intake → open → done.
- Relationships – one case belongs to one vehicle and one client.
This is the very document the agent derived the IndexedDB tables from, and later the Postgres migrations for the backend.
| Field | Type | Required | Source | Notes |
|---|---|---|---|---|
| case_id | uuid | yes | system | PK |
| vehicle_id | uuid (FK) | yes | — | |
| client_id | uuid (FK) | yes | — | Resolved via vehicle |
| insurance_path | enum: pzp_victim / hp_comprehensive / other | yes (at intake) | D011, D050 | Insurance alone decides it (D050) |
| case_status | enum: pending_intake / start / done | yes | D015 | pending_intake = created at the accident call, waiting for intake · start = vehicle on the bench · done = client can pick up |
| accident_time | timestamp | yes | D006, D050 | First-call time ≈ accident. Inline-editable while open |
| total_damage_eur | decimal(10,2) | yes (at done) | D014 | Entered by Ops when closing |
#Calibration: two builds until the prompt became a compiler
##The harness is the prompt
Once the design system, the code rules, and the roadmap with its specs are defined, the first test build can run. That’s the moment the work becomes tangible. If everything’s right, you’ll see a working app with all the UI elements, interactions, and features from the roadmap: fully clickable and finished-looking, built in a single autonomous run of the coding agent. The only prompt I wrote was “Build the whole app from the roadmap.” The harness is your prompt. That was just the trigger.
##The first run is for calibration
The first run is usually a calibration run. It shows how well you’ve defined the design system, the code rules, and above all the roadmap with its specs. Click through the whole app, watch for deviations from what you intended, and trace every inconsistency back to the harness. If something’s off, chances are you’ll fix it by adding a rule, changing one, or defining what was left undefined. The same goes for the product specs: when something doesn’t work or doesn’t look the way it should, you’re usually one or two rule changes in the harness away from fixing it.
For me, it was small stuff: important components I’d forgotten to add to the catalog and the general rules (toasts, for example), a specific sidebar layout for different personas, or details of how form pages are put together.
##Master harness: carry back what the agent figured out
Calibration runs have one more purpose. In the first build you see the whole app for real for the first time, and while building it, the agent also filled the gaps in your harness: it added components that weren’t in the catalog, wrote extra rules, made decisions. Some of them matter. So as the architect and orchestrator, after every calibration run you have to review what the agent added to the harness and the design system. Carry everything important or forgotten back into the starting state of the coding harness, the master version that every new build starts from. That way the harness isn’t calibrated only by your fixes, but also by what the agent figured out on its own.
Then you throw the whole app away and start the run again. It costs you almost nothing, just a bit of time and some tokens from your subscription. And you’ll see for yourself: if your rules are well defined, the app will be 95 percent the same every time.
One of these builds, frontend only, took four hours. I recommend scheduling it as a background task and going to the gym or heading outside in the meantime.
##Iterating with the client
Congratulations, you now have a system where the input is a spec, the output is an app, and the app is disposable: throwing it away is cheaper than fixing it. Now you can go to the client, play with the app, ask for feedback, and iterate until you’re both happy. On this project I ran six frontend builds: the first two were for calibration, and the next four brought in major feedback from the client.
The real magic is that I didn’t implement the feedback by vibe coding right in the code, but at the level of the harness and the specs. I threw away the old app and generated a new one with the feedback already in it, without breaking any existing code.
In practice it looked like this: the client clicked through the app and found that the tables were missing column sorting. I added a rule to the design system: every table must have interactive column sorting. The next build already had it.
Table sorting is a deliberately trivial example. Client feedback usually goes much further: a differently composed flow of screens, a different split of work between roles, a whole chunk of functionality nobody thought of in the workshops. But the path is always the same. The change goes into the harness and the specs, not the code, and the next build already contains it. And this is the biggest value of the whole system: big feedback costs as little as small feedback. In traditional development, the cost of a change is proportional to how much code already exists. Here it’s proportional to how many lines get added to the harness. The next chapter shows that this holds even for a change that breaks the data model.
#When a new module breaks the design: change without pain
One of the worst nightmares on a greenfield project: the product is almost done and you find out you have to rework part of it. It happened to me too. Not because the client wanted a new feature, but because the project had two phases, and while planning the first one I didn’t think through how the second phase’s module would fit into the solution. New scope arrived and broke the app’s current design.
##What happened
Phase two added the fleet module: a portal where a business customer manages their cars, appointments, and service history themselves. The vehicle, the key entity of the whole system, suddenly needed to belong to two worlds at once, internal operations and the fleet, and the phase-one data model hadn’t accounted for that. Some relationships had to be reworked. And the new module brought a whole new world of features into the admin, so the way the screens were split up stopped making sense. Exactly the state you know from ordinary products: a new feature gets added without refactoring the existing ones, and the interface starts to feel chaotic.
##A fix at the architect’s level, not in the code
In traditional development, this change would cost a week or two: someone has to map everything it touches and redo the data model, the migrations, the screens, and everything built on them. We did none of that in the code. The whole fix happened in the harness, at the highest level, where the architect decides:
- Impact assessment. The agent went through the roadmap, all the specs, and the data model and mapped what the new module would break and which specs it would touch.
- Change proposal. Together we designed a new core for the model: the vehicle as a shared entity, with two modules on top of it, internal operations and fleets, that are unaware of each other. The agent carried the change through the whole roadmap and all affected specs at once, so nothing we’d already carefully defined was left inconsistent.
- UI cleanup. The admin navigation was regrouped into new sections so the new module would fit in, not be bolted on.
All in one day. The next day we threw the old app away, and the new build already had the fleet portal in it.
This isn’t a story about a design flaw and its fix. It’s a feature of the system. A scope change that means weeks of pain anywhere else is something we can afford here, because we keep the app in its definition, not in the code. We’re architects: we make the change at the highest level and the code gets recompiled from it. It also helped that this happened in the frontend phase. In my process, the backend is always written at the very end. We had no code to be precious about. We had a definition we fixed, and an app we threw away and had rebuilt.
#Let’s add the backend and start the production run
##Switching the harness to a backend
Once the whole app design is done, every feature is in, the client is happy, and nothing’s missing, it’s time to switch the coding harness from “frontend only” mode to a full app with a backend and production deployment. This is a key checkpoint. It has to be well defined so that a single autonomous run can handle the whole app: the frontend, the backend, and the testing. For the client, nothing changes on the surface: the app looks the same as the one they clicked through. Only the waterline dropped, and everything else gets built below it.
That means going through our coding harness in the docs/ folder and redefining the files that describe the architecture, the boundaries, and the other rules, from a local app to an app with a backend. I write backends in Laravel.
It’s the backend framework with the most complete ecosystem I know: one team builds both the core and official packages for everything a production app needs. Authentication and accounts (Fortify), queues and scheduled jobs (Horizon), emails and notifications, files, monitoring (Pulse, Nightwatch), testing (Pest), a connection to React with no API layer (Inertia), and hosting where the whole thing deploys with a single push (Laravel Cloud). For an agent, that’s ideal: every problem has one canonical, well-documented solution, so there’s no room to improvise. What’s left for me is defining the rules: how to serve the frontend, how to handle authentication and security, how to test the app.
##What survived from the local app
Nothing. Nothing was migrated. The frontend was written from scratch together with the backend, in one run, from the same specs. The only thing that survived from the local version was the design system: the components moved into the new project as a ready-made kit. The IndexedDB schema became Postgres migrations, and the hooks became controllers that send data to the pages via Inertia props. Every build, local and production, lived on its own git branch. Old ones were never edited. We just created a new one. The design phase wasn’t there to produce code we’d keep. Its job was to validate the app design and the data model with the client before we touched the backend and the final handover of the project. The code was a by-product, and it was disposable.
##Two new documents
Compared to the frontend phase, the harness gained two new documents:
SERVICES.md– A binding “problem → package” map: every need the app has (authentication, queues, emails, PDFs, search, files) is assigned one official Laravel package or the framework itself. It’s the first-party packages rule from the stack section, written down so the agent doesn’t have to interpret it on its own.USER_JOURNEYS.md– A catalog of every user journey across roles and portals, each with steps and an expected result. It’s a checklist for clicking through the whole app in the browser, role by role. I’ll show how it was used at the end of the chapter.
##Testing strategy and the verification gate
The testing strategy could easily have its own document in the harness. I put it as a general rule straight into CLAUDE.md, as the only exception to the rule that CLAUDE.md is just an index: it’s too important to make the agent go looking for it. The rule is short: every feature ships with tests in Pest, at minimum the happy path and authorization for every role. In an app with four roles, “who can see and do what” is the most common source of bugs, so this part is mandatory.
Then there’s the verification gate, six commands that have to pass before anything is declared done:
- tests on SQLite, for speed,
- the same tests again on Postgres, to match production,
- backend code formatting,
- a TypeScript type check,
- lint,
- the production build.
And one sentence that turned out to be the most important: the orchestrator runs the gate itself and never takes a subagent’s word that “tests passed.” I’ll explain who the orchestrator is in a moment.
##The orchestrator and two audits
The production run starts with the same prompt as every run before it: “Build the whole app from the roadmap.” The difference is how much the agent has to hold at once: rules for the frontend, backend, UI, UX, tests, and deployment, ten hours straight. For big runs like this, what’s worked for me is not launching a single agent, but an orchestrator that oversees the whole implementation process. For each epic in the roadmap, it launches implementation agents, waits for their reports, checks their work, launches testing and audit agents, and only when the epic is green and committed does it move on to the next one. It writes no code itself. It coordinates and validates.
There are two audits because we know agents don’t always follow the rules a hundred percent. You already know the design audit from the design system chapter: it checks whether the screens stick to the design contract, and whether every new component has a demo in the Showcase and a row in the catalog. The coding audit is new and does the same for the code: it goes through the epic’s frontend and backend and compares them against the rules in our coding harness. Both run as subagents after each epic is finished, the findings get fixed and committed, and only then does the run continue. That’s how you catch most of the cases where the rules aren’t followed. How do you create an audit skill like that? Either you write it by hand, or you let the agent analyze your coding harness and it writes the skill for you. The same goes for the testing strategy.
Or you don’t write anything at all and use existing skills from the framework’s creators. Laravel now publishes its own agent skills: an agent that cleans up freshly written backend code, plus skills for Laravel Cloud and Nightwatch, which cover exactly the deployment and monitoring from my stack. You install them in Claude Code as a plugin, and you’ve got a ready-made layer that guards the framework’s conventions.
##Ten hours without a single question
The production run took ten hours. It was one session running autonomously on my computer, with no interruptions and not a single question from the orchestrator: everything it needed to know was in the harness. The result: 30 migrations, 19 models, 33 controllers, 64 pages, 947 tests with 7,468 assertions, and four user roles and interfaces. A respectable scope for an app.
I didn’t go in blind. We had six frontend runs behind us, so we knew the harness held. Before the live production run, there were two more backend test runs and one full-scope trial production run, just to see where the agents’ limits were at this scale and whether they could handle an autonomous task this long. It wasn’t an experiment but a calculated choice: from previous projects, I had plenty of experience with what agents can pull off. The test runs used Opus 4.8 first, then Fable 5. The difference in code quality wasn’t very noticeable. Both models followed the code rules well. All in all, the app was built ten times, and not once did anything crash or restart. No extra agent costs: the entire run was covered by my regular Anthropic subscription at €200 a month. The production run fit within the subscription’s five-hour usage-limit windows, if barely: when it rolled over from one window to the next, usage was at 99 percent, with the last minute ticking down.
##Audits in action
The audits earned their keep. The coding audit caught, for example, a broken convention for writing backend controllers. The design audit found a page where the agent had invented its own header instead of using the existing PageHeader component, plus new components with no row in the catalog. The findings got fixed and committed, and the run moved on.
##Tests and fifty clicked-through journeys
The agent wrote tests along with every feature, just as the rule in CLAUDE.md requires. The number of tests on its own says little. The number of assertions says how many things each test actually checks, and just under eight per test means the tests aren’t just checking that the page didn’t crash.
But tests can be written badly too, and an agent testing its own code tends to test what it wrote, not what it should have written. That’s why, after the production run finished, we started one more separate run: the agent got USER_JOURNEYS.md, opened a browser, and actually clicked through the app, role by role. That document wasn’t created after the run, but before it: the agent generated it from the roadmap and specs, I reviewed it, and from then on it mirrors everything you can do in the app, role by role. On top of that, the agent was instructed to act like a misbehaving user: try to break the app, enter nonsense, click where it shouldn’t. The same technique comes in handy later in the security review.
It logged in, walked through every flow, opened the downloaded attachments, and checked the sent emails. It fixed nothing, just logged every inconsistency. The result after fifty journeys: no blockers, no functional bugs in the app’s logic. It found two minor bugs, two inconsistencies in the demo data, and one place where the docs didn’t match the code. We fixed all of it afterward.
| Step | Action | Expected |
|---|---|---|
| 1 | Open /admin/fleets | List of portal customers (Customer): companies and individuals alike |
| 2 | Switch the Firmy / Jednotlivci (companies / individuals, B2C) filter | The list filters by customer type |
| 3 | Add a company | The company is created in the invited state; a flash shows the activation link |
| 4 | Open the company detail | The company's vehicles, service records and requests on one page |
| 5 | Add a vehicle: first a free license plate (ŠPZ), then a taken one | The free one goes through; the taken one shows a transfer checkbox and saves only after confirmation (confirm_transfer) |
| 6 | Edit Termíny (due dates): date + reminder | The edit modal saves the date and toggles the reminder; a due date without a date stays „nezadané“ („not set“) |
| 7 | Add a manual service record | The record appears in the vehicle's História (history) |
| 8 | Send the activation | The modal confirms it was sent and shows a one-time activation link, usable in PUB-5 |
| Step | Action | Expected |
|---|---|---|
| 1 | Log in and open /portal | Overview of the single vehicle: the next due date and the „Objednať servis“ („Book a service“) CTA; with no vehicle, an EmptyState shows |
| 2 | Open the vehicle detail | Due dates with a status (green / orange / red) and the service history; the history shows only when the customer ↔ vehicle link is verified |
| 3 | Book a service: type, two preferred dates, fault description, Mobilita (mobility) | The request is sent; the platform stores no exact time, only preferences |
| 4 | Open /portal/bookings | The new request is in the „Čaká“ („Pending“) state |
| 5 | Open its detail and cancel it | requested → cancelled; it disappears from the „Čaká“ filter and shows up under „Zrušené“ („Cancelled“) |
| 6 | Try to open another customer's vehicle detail via a direct URL | They can't reach the detail — the customer ↔ vehicle link isn't verified, so the portal sends them back to its list |
| 7 | Log out and open /portal/bookings via a direct link | Redirect to login; after logging in they land on their requests |
##Security review
In this process, security isn’t a separate discipline at the end, but the result of two decisions: which framework you pick and what you write into its rules. Laravel has most of the defenses built in, and the agent’s job is not to switch them off or get around them. Eloquent sends every query as a prepared statement with bound parameters, so SQL injection has no way through. Every write request needs a CSRF token, and Inertia sends it on its own. React escapes output, so XSS would need a deliberate bypass to get through. $fillable guards against mass assignment, a Form Request validates every input, and authorization lives in policies on the server – not in hidden buttons on the frontend. Fortify hashes passwords, rate-limits login attempts, and handles resets. Documents sit in private object storage, accessible only through an expiring signed URL. I didn’t invent any of it. I just wrote it into the harness as mandatory, so the agent couldn’t get around it, not even by accident.
The second layer was the agent as attacker. We ran the same misbehaving-user technique from the browser run once more, with the Fable 5 model and the reverse brief: break the app. It went through the known classes of attacks on a web app –
- log in as a business customer and use guessed IDs and direct URLs to reach someone else’s case, vehicle, or fleet,
- from a customer account, call admin routes and actions it has no right to,
- slip extra fields into forms, SQL injection into filters and search, XSS payloads into text fields that get echoed back into the interface,
- download someone else’s file from private storage via a tampered URL, reuse a session after logout, brute-force the login.
Result: nothing critical. Not a single attempt got where it shouldn’t have. It’s not luck, and it’s not the model’s doing: with four roles, “who can see and do what” is the most common source of bugs, which is why this process guards it with three layers at once – mandatory authorization tests for every feature, policies as a rule in the harness, and finally an agent that tries to get around it all.
#Was the app done after the production run?
Functionally, yes. Every flow passed, every role got where it was supposed to. Handover-ready? Not yet: there were small things left that my trained eye would catch, not a test.
##Bugs that don’t deserve a rebuild
When I reviewed the code and tested the app, the production run turned out better than I’d expected. Of course, a few small bugs turned up. Looking back, they were mostly missing guidelines for the technical implementation, not logic errors in the code. For example, storing files on the local disk instead of external storage: after the next deployment, all the files would have been deleted.
The other bugs were in the same category. I’m picking them straight from the commit messages in git:
- A “back” button in a form that accidentally submitted the form because it had no type set.
- One login card with no inner padding, while all the others had it.
- A modal that couldn’t be scrolled on short screens.
- A percentage field that didn’t accept a decimal comma, only a period.
- A text template that added a period after a value that already ended with one.
- Ambiguous file naming.
Cosmetics, setup, small conventions. Not one of them was architectural, and that’s an important guideline for this whole approach. A bug deserves a rebuild with updated rules when it carries a fundamental problem in the architecture and, without a rule, would repeat on every screen that follows.
An example from our context: suppose the agent handled permissions on the frontend, hiding buttons and menu items by role, instead of on the server through policies. Every screen would inherit it, every one of them could be bypassed with a direct URL, and a fix in the code would mean going through all 64 pages. That bug calls for a rule and a rebuild. A bug that’s one-off and local doesn’t deserve a rebuild. It would be an expensive answer to a cheap problem.
##You point at the bug, not the fix
So how did they get fixed? Not with vibe coding, where you dictate to the agent what to rewrite and where. The agent got the list of findings, just as the browser run and my own review had logged them, and worked through them autonomously: for each one it proposed a fix itself, implemented it by the harness rules, and ran the verification gate. I checked every fix before we saved it. You point at the bug, not the fix.
Personally, I think code today is better written than it was at the peak of my active engineering career, when I was building apps every day.
##The last human touch
The agents stuck to the design contract from the rules in the repo, and the UI felt solid to me. But as a retired UI/UX designer and a perfectionist, I spent roughly two more days polishing the UI and the flow, so that even the smallest detail and the smallest action in the app would feel pleasant and natural. For this part I have no carefully designed process, just good old vibe coding: Claude, do this like this, do that like that. It’s polishing small details, not vibe coding new features. A bit of the human touch the production version needed. It was the last human touch before handing over the project: a polish, not a rebuild.
One autonomous run gets you an app that works, holds the design, and can be handed over after two days of polishing. That quality isn’t made in the run, but before it: in a harness with clear rules for how the app should look, what interaction pattern it has, and how it should be built. The last few percent are just a question of your eye for detail. And those get polished in the code, not in another run.
#What was left after the production run
Code. For now. There’s a point in my process where I switch over: the app becomes the single source of truth, and from then on we iterate on it, directly, if development is ongoing and the client keeps bringing new business requirements. The specs and the harness stay as documentation of why the app looks the way it does, but nothing gets compiled from them anymore.
##Six weeks instead of twelve
The whole execution phase of the project took six weeks. Before the era of coding agents, in my experience, the same scope took twelve weeks and needed a product manager, a designer, and two to three engineers. Today one person with a foot in product, design, and engineering can orchestrate it. And on top of that you get something we didn’t have before AI: the client has the app in their hands early in the solution design, feedback comes back in hours, and every build is the whole product – not a Figma prototype, not a piece of the app, not a stripped-down MVP. You don’t build code anymore. You have just two jobs: architect the right solution and orchestrate the agents that build it.
#What was hard
##Managing the specs
The hardest part of the whole process is managing the specs for every feature in the roadmap. When you have, say, a hundred of them and you work with them all day, it’s genuinely exhausting, and you have to decide for yourself whether the trade-off is worth it. This is where you need the full power of agents: let them help you manage the specs. And keep them lean and simple so they don’t grind you down later: today’s agents aren’t dumb, and even from minimal information they’ll understand what you need. Always keep the key information at the harness level, so you put as few duplicate instructions into the specs as possible.
##A complex UX flow as a prototype
The second hard thing is describing a more complex UX flow in a spec, say a ten-step process the user has to go through. Describing it well enough that the agent replicates it the way you picture it is hard. Especially when you recompile the app again and again in the iterative phase. What worked for me was to take the design system and iterate my way (yes, by vibe coding) to a standalone prototype with a UI that does exactly that tricky flow. I saved it in the harness under docs/prototypes, and the spec for that feature contains only a reference: the agent looks at it and implements it in the production code cleanly and by the production rules. That way you get the best of both worlds, vibe coding and spec as code.
#One level up
##The prompt, wait, check loop
I’m noticing that a lot of engineers are losing the joy they used to get from their work because of AI agents. Programming was their everyday work, and AI is taking the joy out of it. An engineer’s typical day used to go like this. You get up in the morning, make coffee, open your IDE, and dive into a problem for eight hours: you break it into parts, build a solution, and by evening it works. You created something, solved something, and that was your reward. Today the same person sits all day in a prompt, wait, check, prompt loop. They make one small decision after another, the agent does the work in between, and by evening they’re exhausted from deciding, with no sense of having built anything.
##The work has changed, and so has the reward
I see it differently: you have to go one level up. We don’t build the app at the code level anymore. We build the harness: the knowledge that steers the agent, the design system, the tests, and the audit, so the agent builds the app from scratch in one autonomous run. It’s different work, and judgment moves higher: what technologies to build with, what the UX and UI should look like, how the system should behave when something is missing. And you can play and iterate with it just like with code. The reward just lives somewhere else. Not a feature that works, but a system that works. And what comes out of it is three times bigger than what we’ve built so far.
The question is no longer how to build a piece of code well. The question is how far to raise your ambitions. To keep that reward from disappearing, the scope of what we can do has to grow. Once, we hunted game, then sowed fields, then wrote code that automated one thing. Today we can design a whole system that grows and changes with what the world around it wants from it. That’s in our hands now: to design it and orchestrate it.
#Where this is heading
##The harness is disposable too
This article is more about a mental model than an exact process. It shows that it can be done this way, not that it has to be done exactly this way. Don’t copy my harness documents. Take inspiration and try what fits you: every person and every project has their own ways of working and is at their own stage. And the harness you build today won’t be relevant in six months. It’s nonsense to believe it will.
What’s true for the app is true for the harness: it’s disposable. Only one thing has to stay: the willingness to throw away the old mental model when the world changes, and build a new one. Today, AI itself will build the harness for you. What it won’t build for you is flexible thinking and the courage to experiment.
##The product builder in the field
Here’s where I’m heading. Building greenfield projects fast, where a design flaw isn’t a disaster but a cheap fix. Where a feature added mid-flight isn’t a curse, it’s a feature. Where I, as a product builder, spend more time in the field talking with the customer and the users than sitting at the drawing board and in front of coding screens. And where the vision of a solution can be realized a lot sooner and more cheaply than was possible before. Not a stripped-down MVP, but a finished product.
##Next experiment: fixed data, replaceable app
And one thing I want to try next time. Is the final production code really the source of truth, or is it still the product spec? I’m toying with the idea of separating out the production database as a static element and leaving the app as a dynamic element that can be recompiled again and again, even through further big iterations. Whether you add a feature or remove one, you always recompile the app. Picked up a huge number of users and the app can’t keep up? You recompile the backend with a different architecture, or in a different language or environment. The production data your customers and your company generated is fixed. Everything above it is replaceable. I’ll save that for future experiments, maybe for a future article.
And that’s just one idea. As I write this, a huge number of problems and opportunities come to mind that could be reinvented this way. I’m sure that while reading, you came up with at least ten things for another ten articles. I’m not afraid that AI will leave us with nothing to do. Quite the opposite: there’ll be a lot more work. It’ll just be different.
If you’re building something similar, like your own harness, a spec as the source of truth, or a disposable app, get in touch. I’d be glad to compare notes.
AI-era evangelist. I show what one person can do today: I build generative AI solutions and share what I learn along the way. Get inspired, take what’s useful, and go build. Working on something similar? Message me.