definition compiler ×9

Recompile, don’t refactor. How I build production software with AI: not the app, but the system that compiles it

Agentic Engineering·September 22, 2026·
Article summaryAI
  1. Don’t build the app. Build the system that compiles it. The real prompt is the coding harness: code rules, a design system, a roadmap with specs, and a data model. The only prompt I wrote was “Build the whole app from the roadmap.”
  2. The app is disposable. Feedback and scope changes go into the harness, not the code, and the next build already has them baked in. The app was regenerated from scratch ten times, and I didn’t touch the code once.
  3. Big feedback costs as little as small feedback. The cost of a change isn’t proportional to the code that already exists, but to the lines added to the harness. When a phase-two module broke the data model, the fix took a day, and then a new build ran.
  4. The client tests from week one. The first version of the app is always just a frontend, indistinguishable from production. Four iterations with the client validated the product, the scope, and the data model before the backend existed. The code from this phase was a by-product.
  5. A component library instead of a tokens doc. The agent never followed a tokens doc. Vibe coding is only there to find the look. The agent extracted a library from that look and maintains it on its own through the 80 percent rule, derivation recipes, and a design audit.
  6. Ten hours, one prompt, zero questions. The orchestrator runs the agents epic by epic through a verification gate and two audits, and never trusts that “tests passed.” Result: 64 pages, 947 tests, and fifty clicked-through journeys without a blocker.
  7. Go one level up. Six weeks instead of twelve, one person instead of a team. Only a bug that’s missing a rule deserves a rebuild. The reward is no longer a piece of code that works, but a system that generates it right. And this harness is disposable too.

AI wrote this summary from the full article. The author checked it.

I don’t write code anymore. I change the product definition at the top, and coding agents recompile it into clean application code in a few hours, again and again, with no technical debt. This is how I build production apps for clients, from scratch, built entirely with AI.

Four days before the deadline, all I had was a system that compiles. No app. Before agentic engineering, by this point you’d already be a few weeks into testing and debugging the app. Here the design was done, but the app’s code didn’t exist. I kicked off a ten-hour production build with AI, and by evening we had a finished production app for the client. It was quite an experience: living on the bleeding edge, where you trust the system you built, not the code you see.

For over seventeen years, I’ve been architecting, designing, and coding apps with business owners. After four years building apps in a team as a product manager, I got back into a delivery role: I could once again own the end-to-end process myself, think hard about AI, and reinvent how an app gets made. This article is about how I design and build greenfield production apps for clients today, built entirely with AI: one client project, from nine workshops to a ten-hour production run.

#Three things that changed

I didn’t go the vibe coding route, where you build the app one prompt at a time and every change is another patch. Without an architecture decided at the start, at some point it turns into chaos and technical debt that stops your project in its tracks. Instead, I built a system that reinvents three things software development has stood on for seventeen years. In this process, they’re game changers.

+1 +1 +1 code = source of truth 30 changes later: a pile of accidents, chaos vs harness = source of truth app always rebuilt from scratch +1 +1 +1 code = source of truth 30 changes later: a pile of accidents, chaos vs harness = source of truth the app is always rebuilt from scratch
  • The app is disposable. I don’t write any code. I change the product definition at the top, and in a few hours the agents recompile it into clean code from scratch, with no technical debt and every change applied. On this project, the app was built from scratch ten times, and I didn’t touch the code once. A scope change that broke the data model, and would have cost a week of refactoring in code, cost a day and one build here.

  • The client tests from week one, not at the end. The biggest risk of traditional development: the product is almost done, and only during testing does the client find out that something is missing or badly designed. Here, from week one, the client had an app in hand that was indistinguishable from production, in almost full scope. They gave feedback and saw a new version the next day. Four times. By the time the real production app was built, the product, the scope, and the data model had already been validated.

  • One person, one prompt, ten hours. One autonomous agentic run on my computer built the whole production app, from one prompt, without a single question. Executing the whole project took six weeks. Before the era of agents, the same scope would have taken twelve weeks and needed a product manager, a designer, and two to three engineers.

About the app: the client, Marco CAR, an automotive company in Slovakia with fewer than 50 employees, commissioned an app from me to run its internal operations, plus a customer-facing module where business customers manage their own fleets. Repair cases, vehicles, appointments, service history, notifications. Four user roles, each with its own interface, eight modules, two phases.

The rest of this article is how I built it all.

#Before the first build

##Nine workshops

Before the build itself came nine workshops with the client, ~90 minutes each. The goal was to understand the requirements, explore the company’s business processes, and clearly name the problems its custom software should solve. We worked on two tracks: designing new internal operations and the software to run them, and a customer-facing module of the same software, because the company cares about its customers’ long-term satisfaction. After the workshops, I knew what to build.

From the very first workshop, AI helped me store the information and give it a clear structure. Throughout the project, that structure was the main source of truth for decisions about the system’s design.

##Two harnesses

This is a good place to introduce two terms that will keep coming up in this article. The project has two harnesses, and each one does something different.

The knowledge harness is my own system for organizing the information from all the workshops with the client and synthesizing it into the final outputs: the roadmap and the feature specs. It’s also the project’s operating system: meeting notes, decisions along with their reasons, a map of the people on the client’s side, all the communication. From that same state, it puts together my meeting materials, an email to the client, a proposal, and a contract, and it answers the question “why did we decide on this?” with a link to the specific note. How it works on the inside I’ll save for a separate article.

The coding harness is a set of text documents in the app’s repository that define the rules for the coding agent: how to write the frontend, how to write the backend, what the architecture is, how to test, and how to deploy. It’s the prompt the agent gets instead of a brief. The rest of this article is about it. From here on, when I write just “harness,” I mean this coding harness.

The relationship between them runs one way: the workshops go into the knowledge harness; out of it come the roadmap, the specs, and the data model; and those go into the coding harness as the brief for what to build.

9 workshops knowledge harness scope roadmap · specs · model coding harness compiler app disposable artifact 9 workshops knowledge harness scope roadmap · specs · model coding harness compiler app disposable artifact

##Tech stack

Let’s look at my tech stack:

  • Agent – Claude Code, extended with my own skills for synthesis and management in each phase of the project: discovery, roadmap, spec, build, design/code audit.
  • Frontend – React 19 with TypeScript and Tailwind v4, built with Vite, no UI library: every component is custom, and I’ll explain why when we get to the design system. In the design phase it ran on its own, without a backend, with IndexedDB as a local database.
  • Backend – Laravel 13. The frontend sits on top of it through Inertia, so the app is a monolith with no API layer: a controller sends data to the page as props, and a form sends it back.
  • Communication and integrations – Emails through Laravel Mail and Notifications (account activation, monthly reports, reminders), via Resend in production. SMS through BulkGate, with no package, just a thin custom channel over the HTTP client. PDF through dompdf to generate internal documents. The scheduler for month-end closing and appointment reminders. Integration with SoftApp, the service shop’s inventory and invoicing system, with service records matched by license plate.
  • Infrastructure – PostgreSQL, Redis with Horizon for queues, Fortify for authentication, private object storage for invoices and statements, monitoring through Pulse and Nightwatch, tests in Pest, hosting on Laravel Cloud.

One rule in the harness is worth mentioning: every problem is solved with a first-party Laravel package, and nothing else gets installed without my approval and an entry in the decision log. That leaves the agent no room to improvise with libraries.

#By the numbers

#How do you generate an app ten times so it looks and works the same every time?

And on top of that, with the client’s feedback built in each time. Without an answer to this question, you can’t iterate like this. You don’t want to hand the client the app for the tenth time and have it look different every time.

The answer is the coding harness, our compiler. Everything that has to stay stable between builds is defined at its level. The design system: how the app looks, what interaction patterns and rules it has, and how to derive new components that aren’t in it yet. The code rules: exactly how to write the frontend, how to write the backend, what the architecture is, what the testing strategy is, how tests are written, and how deployment works. And finally, the backbone of the app: the full scope of features and a spec for each one. Everything that matters is defined once, at the harness level.

“Build the whole app from the roadmap.” 7 words rules FE · BE · design contract tests · smoke · deploy strategy rules & quality audit data model · entities roadmap + 64 specs harness = the real prompt · 196,000 words app = artifact · 10 h “Build the whole app from the roadmap.” 7 words rules FE · BE · design contract tests · smoke · deploy strategy rules & quality audit data model · entities roadmap + 64 specs harness = the real prompt · 196,000 words app = artifact · 10 h

The next three chapters walk through these layers: the design system, the frontend rules, and the roadmap with its specs.

#The app’s design system

I wanted to reinvent how a brand-new product comes to life, from the first idea to production. So let’s start at the beginning: how I created a design system without opening Figma and starting to draw screens.

##Vibe coding to find the look

I already had a picture in my head: what dashboard the company needed and what visual language it should speak. I focused on functional design, not aesthetics, and on that basis I started to vibe-code a dashboard for the business owner. (The examples below are real screens and the component catalog from this project, published with the client’s consent. It’s only part of the app, not the whole product.)

This is exactly where vibe coding fits in my process: finding the look, not building the product. It was nothing complicated, and you don’t need to think about best practices for the frontend stack. A new folder, Claude Code open, and a prompt to create a basic React project with the owner’s dashboard page in it, using the colors, fonts, and design direction I gave it. I wanted an admin dashboard with a navigation sidebar on the left and typical admin content on the right: tables, tabs, search, a few stats.

From a rough sketch to the dashboard’s final look — component by component, then styled step by step. The brand color only arrives at the end.

I then iterated on the first version and polished it several times until I liked how it looked. The basic dashboard was maybe 15 to 20 prompts. And it was interactive: the components responded to clicks. It wasn’t a static screen.

That’s how I ended up with four dashboard variants in total, each with a different kind of look. I picked the one I liked most for the project’s purposes and derived three more pages from it: an entity detail with information, tables, and action buttons; a list of entities in a table; and a form to create a new entity. These four pages contained most of the interaction elements that repeat across the app as its basic building blocks. And I could show them to the client and ask for feedback on both the look and the interaction patterns before a single real feature existed.

Four pages of the app plus a modal: dashboard, list, detail, form. They contain most of the interaction elements.

##From four pages to a UI component library

It was two days of manual work exploring and building the foundation of the app’s design. The next step was the key one: I had Claude extract every UI component from those four vibe-coded pages and turn them into a full-fledged UI component library for the app. To make them visible and testable, I also created a Showcase page where you can browse them all, with descriptions: fonts, colors, whitespace, and what each component is for and how to use it.

The Showcase page: every component in one place, with a description of what it’s for.

The library has one purpose: to make the coding agent build the frontend consistently across the whole app, while I keep full control over how it looks. I also tried a simpler route, a single design.md document with tokens: colors, fonts, paddings. It didn’t work. The agent never stuck to the look, and in every run it derived a different kind of app, never a consistent one. A custom component library, finished before the first build, turned out to be the only reliable way to get the same frontend across ten runs.

design.md tokens in text → 5 variants of one modal vs <DesignSystem /> components in code → one modal, every time design.md tokens in text → 5 variants of one modal vs <DesignSystem /> components in code → one modal, every time

A few numbers, to make it tangible. From those four vibe-coded pages, Claude extracted 46 components. Today, after the production run, the catalog has 83: 63 primitives (buttons, fields, pills, cards, tables, drawer, modal), 4 composites (app shell, sidebar, page header, form card), and 16 domain components like the case timeline or the service history table. I didn’t write a single one of those 37 new ones. The agent derived all of them during the builds.

##Why no third-party UI library

This is the right place to say why the project has no third-party UI library. I’ve never used them. I’ve always preferred to design and code my own components. For me, UI libraries are limiting both aesthetically and functionally: they bring chaos and hidden complexity into the code and push their own opinionated direction on how an interface should look. I like clean, custom solutions. And today it’s no longer a trade-off between speed and cleanliness: a coding agent codes up a component with everything we need, on the spot, with none of the baggage a library would bring along. The deliberate trade-off is accessibility, which ready-made libraries like shadcn/ui handle for you (keyboard, focus, ARIA...), but it’s nothing the agent can’t manage when the project calls for it.

third-party UI library Save deadweight chaos, hidden complexity, their opinionated direction tailor-made component Save exactly what’s needed, agent · on the spot · no baggage third-party UI library Save deadweight chaos, hidden complexity, their opinionated direction tailor-made component Save exactly what’s needed, agent · on the spot · no baggage

##The design contract: how the system is used

Once the components were coded, I wrote a design contract into the harness. It’s not a description of the look. That already lives in the code and on the Showcase page. It’s a manual for how the system is used and where its boundaries are. It has seven parts:

  • Agent contract – six steps in a fixed order that the agent goes through for every screen: pick a page template, pick components by purpose, derive only when nothing fits, never invent tokens, put every new component in the Showcase, stick to the localization.
  • App shell and canvas – one canonical shell for all roles, no top bar, white canvas, sepia only in drawers, fixed sidebar dimensions, full-width content. Three page templates (dashboard, list + detail, centered form) that every new screen is derived from. A drawer only on the dashboard for a quick preview, a modal for one focused question. A single breakpoint where the shell switches to mobile mode.
  • Component selection – an intent → component table: primary action, secondary action, status pill, search, form field, KPI tile, empty state, error. The agent first names what it wants and only then reaches for a component. Plus rules for tables: column order, row actions, server-side sorting.
  • Derivation recipes – eight shapes (button, field, toggle, pill, tile, surface, overlay, composite), and for each one a procedure for what to derive a new variant from and how.
  • Icon contract – size and stroke weight by context, colors, bi-tone active icons in the navigation.
  • Anti-patterns – a list of things that made it into the code at least once, in this project or a sister project: no raw hex, no new accent colors, no orange buttons, no heavy shadows, no font-bold, no third-party UI library, no made-up variants.
  • When this document is wrong – if the rules block a reasonable solution, that’s a signal to escalate, not permission to break them.
Design System — Agent Operating Manual
This document is the contract for any agent (or human) generating UI in this codebase. It does not restate colors, typography, or component APIs — those have a single source of truth elsewhere:
Looking for…Go to…
Color, radii, shadow, typography tokenssrc/index.css (@theme block)
Live, runnable showcase of every primitivesrc/pages/DesignSystem.tsx (route /design-system)
One-line catalog of every component + propsdocs/COMPONENTS.md
Layer boundaries / import rulesdocs/BOUNDARIES.md
Project architecture, hooks, servicesdocs/ARCHITECTURE.md
Routes & navigationdocs/ROUTES.md
Error UX patterndocs/ERRORS.md
This file owns: how to use the system, when to derive new components, and what's forbidden.
1. Agent contract
When you are asked to generate or modify UI, run this loop in order. Do not skip steps.
1.Pick a page template. Open src/pages/DesignSystem.tsx → PageTemplatesSection. Three canonical layouts exist: dashboard, list+detail, centered form. Start from the one whose shape matches the screen.
2.Pick primitives by purpose, not by name. Walk the showcase from top to bottom. For each piece of the screen, find the closest existing component in src/components/ (catalog in docs/COMPONENTS.md). If something fits within ~80%, use it and pass props — don't fork it.
3.Only derive when no primitive fits. If you can't find a match, follow §4 Derivation recipes below before writing anything new. Composing existing primitives is always preferred over building a new one.
4.Never invent tokens. Colors, radii, shadows, motion durations all come from src/index.css. If a value seems missing, you are almost certainly trying to do something the system rejects — re-read §6 Anti-patterns.
5.Showcase what you build. Any new primitive added to src/components/ must also appear in src/pages/DesignSystem.tsx with a Showpiece and get a row in docs/COMPONENTS.md. No exceptions — the showcase is the system's source of truth.
6.Match the locale. App copy is Slovak, informal-direct register, no exclamation marks. Reuse phrasing from the showcase or existing pages.
2. App shell, canvas & locale
The contract in §1 covers individual components. This section covers the things the agent loses control of when it improvises chrome — and where past builds have shipped the wrong answer most often.
2.1 App shell is one canonical component
The chrome that wraps every authenticated screen (sidebar + main canvas + create menu + user footer) lives in one file: src/components/AppShell.tsx. It is composed from Sidebar + NavSection + NavItem + CreateMenu exactly as SidebarShellDemo in src/pages/DesignSystem.tsx shows.
•Pages render only their content inside <Outlet />. Shell concerns — nav, create menu, user footer, canvas color, container padding — belong to AppShell, never to route files.
•Different portals (admin vs zákaznícky, Owner vs Ops) are modes of the same shell, expressed through props or conditional nav arrays. Do not create a parallel shell implementation per portal.
•If SidebarShellDemo lives only inside the showcase, extract it to src/components/AppShell.tsx first and import it from both the showcase and production. Two implementations of the same chrome are never allowed.
•There is no top bar. The showcase templates have none. Page titles belong to PageHeader, not the chrome.
2.2 Where the CreateMenu lives
CreateMenu is the dark "Vytvoriť" button. It belongs inside the sidebar via <Sidebar topSlot={<CreateMenu … />}>, directly below the logo. It is never placed in a topbar, header, or floating in the page body. The showcase demonstrates this in SidebarShellDemo.
2.3 Main canvas color
The main canvas (everywhere outside drawers and modals) is bg-bg-surface — white. Light grey (bg-bg-base) is reserved for drawer shells only — and drawers appear only on the Dashboard (see §2.6). Its purpose there is to make the inner white DetailCards lift visually. Using grey on the main canvas erodes that signal.
The showcase's PageTemplatesSection description states the same rule: the main canvas is white, bg-bg-base belongs to drawer shells.
The palette is neutral: white canvas, light-grey bg-bg-base, ink text (#171717, never pure black), one orange accent (primary) and a muted red used for danger only. The accent-red and accent-lime tokens both resolve to ink — they name a role (structure, toggles/focus), not a hue. Fonts: Ubuntu (sans) and Ubuntu Mono (mono). All values live in src/index.css.
2.4 Sidebar dimensions
Sidebar expanded width: 260px. Collapsed: 76px. These are the values SidebarShellDemo uses. Do not improvise narrower or wider sidebars per page — the rest of the layout grid is calibrated to these numbers.
2.5 Page templates
The showcase's PageTemplatesSection defines three layouts. Every route-level page must match one of them:
•Dashboard — KPI tile row + content cards, sidebar + main canvas, no extra topbar. This is the only screen that opens a Drawer — see §2.6.
•List + detail — table on the main canvas; selecting a row navigates to a full-page detail route (BackLink + DetailPageHeader + a column of DetailCards on the white canvas). The detail is its own page — never a drawer. Used for cases, clients, vehicles.
•Centered form — max-w-[560px] mx-auto, BackLink above PageHeader, single FormCard. Used for invites, settings sub-pages, two-step flows. Wider or left-aligned form pages are a deviation.
If a screen doesn't fit any of the three, stop and surface the conflict before improvising — adding a fourth template is an owner-approval moment, not an in-flight decision.
2.6 Drawers are Dashboard-only
The Drawer (right-aligned overlay, light-grey inner shell) appears on one screen — the Dashboard. It has exactly two jobs there:
•Quick preview — clicking a record in a dashboard widget opens a read-only glance. The drawer's action row carries an open-full-detail control (e.g. „Otvoriť nehodu") that navigates to the full-page detail route. You do not edit inside the drawer — it is a glance, not a workspace.
•Aktivita panel — the activity feed (TimelineFeed), opened from the bell.
Every other screen resolves record detail and activity as a full page, never a drawer:
•A list page (Nehody, Klienti, Flotily) opens a row as a full-page detail route — BackLink + DetailPageHeader + DetailCards on the white canvas.
•Activity reached outside the Dashboard is a full page, not the Aktivita drawer.
If you are building a non-Dashboard screen and reach for Drawer, stop — the system wants a full-page detail (template §2.5) instead.
2.7 Locale
App copy is Slovak, informal-direct register, no exclamation marks. When in doubt about phrasing, lift it from the showcase or an existing page rather than translating freshly. Mixing English UI strings or formal address ("Vy"/"Vám" outside of explicit legal copy) breaks the voice.
2.8 Main content spans the full canvas width
The content area — everything to the right of the sidebar — fills 100% of the remaining width. It is a fluid 1fr column, never a fixed pixel width and never a centered max-w-* wrapper at the canvas level.
•The shell grid is [sidebar] [1fr]. The content column stretches and shrinks with the viewport; only the sidebar carries fixed widths (§2.4).
•Width constraints belong to inner elements, not the canvas. The centered-form template (§2.5) caps its FormCard at max-w-[560px] — but that cap lives on the card, while the canvas underneath still spans the full column.
•Do not wrap a page's content in a fixed-width or max-w-* container to "keep it tidy". Tables, dashboards, and detail pages use the whole width.
The PageTemplatesSection templates already encode this — their main column is 1fr. Copy that, never reintroduce a fixed shell width.
3. Component selection rules
Map the intent → primitive before writing JSX. Use this table as a decision tree:
IntentPrimitive
Hero action on a pagePrimaryCTA (dark fill). Never orange.
Cancel / secondary actionSecondaryButton
Toolbar action with icon + labelGhostButton
Icon-only button (with aria-label)IconButton
Status labelPill (semantic tones) — or StatePill for case lifecycle
Hero search bar (Dashboard)HeroSearch — live result list inside the card
In-table / compact searchTableSearch (inside TablePanel)
Form fieldTextField (mono / prefix / suffix / error all built-in)
Two/three-option choice in a formSegmentField
Yes/no settingSwitch (ink default; orange only for network-level settings)
Card list of options where one is selectedRadioGroup
KPI numbers above a tableStatTile (only one highlight tile per row)
Container with elevationCard
Container without elevationSurface (tints: surface, base, subtle)
Filterable listTablePanel + DataTable + TableRow + TableCell
Empty listEmptyState
Inline noticeAlertBanner (tones: danger, warn, info)
Record detail opened from a list pageFull-page detail route — BackLink + DetailPageHeader + DetailCards. Never a drawer (see §2.6)
Quick record preview or the activity panel — Dashboard onlyDrawer (right-aligned overlay) — see §2.6
Single-question editModal (centered, ~480px)
Numeric ± edit inside a modalStepper + DiffStrip
Inner-page title blockPageHeader + BackLink (list / form pages); DetailPageHeader + BackLink (entity profile pages)
Section divider with accentSectionAccent + SectionLabel, or SectionHeader for the full block
Activity feed grouped by dayTimelineFeed
ID / serial number / installation number / contract number / shortcodewrap in <Mono>
Loading stateSkeleton (shape-matched) for surface fills, Spinner for inline, PrimaryCTA loading for buttons
Failure stateErrorDialog (centered modal — never inline error rows; see docs/ERRORS.md)
Anti-rule. Reaching for a third-party UI library (Headless UI, Radix, MUI, shadcn) is forbidden without explicit owner approval. Every primitive listed above is already implemented locally.
3.1 Table column order
Every DataTable orders its columns by a fixed priority. Do not reorder per screen — a consistent skeleton lets the operator scan any table the same way.
1.ID — the entity's identifier (client ID, contract number), wrapped in <Mono>. Include this column only when the entity has an ID; skip the slot entirely when it doesn't.
2.Entity name — the human label: client name, company name, case title.
3.Labels — status and category Pills / StatePills (case state, type, tags). All label-style columns group here, directly after the name.
4.Everything else — remaining data columns (serial number, installation number, amounts, dates, owner), in whatever order best fits the screen.
5.Actions — the row-action column, always last (§3.2).
So a case table reads: Klient · Typ · Stav · ŠPZ · Škoda · Akcie. TableSection in src/pages/DesignSystem.tsx is the reference — mirror it.
3.2 Table row actions are icon-only
Row-action buttons never carry a text label. Each action is a single clickable symbol — a ghost IconButton with a required aria-label (and a matching title for the hover tooltip). This holds for every row action: open, edit, delete, download.
•Use IconButton variant="ghost" size="sm"; the action column is the last one and its column field is sortable: false.
•A text-labelled button (PrimaryCTA, SecondaryButton, or GhostButton with children) inside a table row is a deviation — see §6.
•Page-level actions that do need a label live in the PageHeader or the TablePanel controls row, never in the row itself.
3.3 Cross-link client / vehicle / case references
Anywhere an entity — a client, a vehicle, or a case — is named, that name is a navigable link to the entity's detail route, not plain text. This applies everywhere the reference appears: table cells, DetailCard values, drawer previews, timeline items, activity feeds, modal copy.
•Render the reference as a link to the entity's full-page detail route (List + detail template, §2.5). The operator can always pivot from one record to a related one in a single click.
•In a table, both the ID column and the name column link to the same detail route.
•Plain, non-clickable text for a client/vehicle/case name is a deviation. If a reference genuinely has no detail route to point at, that is a gap to surface — not a reason to drop the link.
•If a shared link primitive is needed for this, derive it per §4 — keep the styling calm (inherit text color; underline or color shift on hover only).
4. Derivation recipes
If §3 has no fit, follow the closest recipe below. Every new component:
•Goes in src/components/<Name>.tsx, named export, functional only, no any.
•Uses only token utility classes from src/index.css (bg-bg-surface, text-text-primary, border-border-strong, shadow-card, …). No raw hex anywhere.
•Inherits API conventions from its closest sibling primitive (see recipes).
•Gets a Showpiece in DesignSystem.tsx and a row in docs/COMPONENTS.md in the same PR.
4.1 New button shape
Copy the API surface of PrimaryCTA: { children, icon?, trailingIcon?, size?, loading?, ...buttonProps }. Pass ...rest through to <button> so consumers retain disabled, onClick, aria-*. Pick the existing fill family — dark, light-with-strong-border, or ghost — don't invent a fourth.
4.2 New form field
Mirror TextField's outer envelope: label row (12px medium text-text-secondary, ink * for required) → bordered input shell (border-border-strong resting, border-accent-lime — ink — on focus) → optional hint or errorMessage (11.8px, text-text-muted or text-danger-text). All form fields must accept label, required, hint, errorMessage. If the control is intrinsically wide (textarea, file dropper), keep the same envelope but swap the inner element.
4.3 New toggle / boolean
Choose between Switch (continuous setting) and Checkbox (acknowledgement / multi-select). If you must introduce a new toggle pattern (e.g., tri-state), still keep the props: { checked / value, onChange, label, hint?, disabled?, tone? }. Tones are lime (default, renders as ink) and orange (network-level) — do not invent a new tone.
4.4 New pill / chip
Use Pill first with one of its tones. Only create a new chip component when the shape differs (e.g., a chip with a leading avatar, or a removable chip with an ×). Reuse the same radii (rounded-full for chips, rounded-md for square pills) and the soft/strong background pair from the existing tones.
4.5 New stat / tile
Use StatTile with a custom tile payload before forking it. If you need a different shape (e.g., a sparkline tile), keep StatTile's outer card shell (border, radius, padding, hover lift) and only change the inner block. Maintain the one-highlight-tile-per-row rule.
4.6 New surface
Card (elevated) vs Surface (flat). Do not create a third surface treatment. If you need a tinted surface, use Surface with tint="subtle" | "base".
4.7 New overlay
Two kinds exist and that is final:
•Drawer — right-aligned, full-height. Dashboard-only (see §2.6): a quick read-only record preview, or the Aktivita panel. Always uses Drawer + DrawerHeader + DrawerBody. Record detail on any non-Dashboard screen is a full-page route, not a drawer.
•Modal — centered, fixed width, single focused question. Always uses Modal.
Anything else (popovers, command palettes) requires owner approval before you start.
Sanctioned exception — success toast. The bottom-right success toast (ToastProvider + useToast, src/components/Toast.tsx) is an owner-approved non-blocking confirmation surface for completed actions (client registered, case opened, invoice uploaded). It is success-only: errors still go through ErrorDialog per docs/ERRORS.md. Do not stretch the toast into warnings, info banners, or destructive confirmations — those have their own primitives (AlertBanner, Modal).
<!-- Trimmed for the article — one more sanctioned exception stands here in the real doc (and in yours). -->
4.8 New composite (multi-primitive)
Composites that bind primitives to a domain (InvoiceUploadModal, BugReportModal) live under src/components/ only if they are reused across ≥2 pages. Single-use compositions stay inside their pages/<feature>/ folder. Always compose from primitives — never duplicate token classes.
5. Icon contract
The icon library is lucide-react. There are exactly three rules.
5.1 Size & stroke
ContextSizeStroke
Nav rail / sidebar item202
Inline with text / list rows161.75
Ghost button leading icon141.75
Hero search / page-level181.75
Big empty/error state22–241.5
Use these exact pairs. Avoid strokeWidth={1} (too thin) and strokeWidth={2.25+} (too punchy).
5.2 Color
Icons follow text color unless they carry semantics. Defaults:
•Calm / structural: text-text-muted or text-text-secondary.
•Inside a tone container (alert/pill): inherit the container's text-*-text.
•Inside a PrimaryCTA / cta-dark: white (text-text-inverse).
Never set an icon's stroke directly via style. If you need a custom color, wrap it in a span with text-* — the icon inherits.
5.3 Bi-tone active treatment (sidebar nav)
When an icon needs the brand bi-tone treatment (ink structure, orange accent on a secondary path), wrap it in .mc-bitone[data-active="true|false"]. See src/pages/DesignSystem.tsx → BiToneIconsSection for the live demo.
To add a new icon to the bi-tone set, append a CSS rule to src/index.css under the mc-bitone block:
.mc-bitone[data-active="true"] svg.lucide-<icon-name> > *:nth-child(N) { stroke: var(--color-primary); }
Steps:
1.Inspect the icon's lucide SVG to count children (path, rect, circle, …). Lucide exposes them in source order.
2.Decide which child is the accent — typically a secondary, visually smaller element: a dot, a notch, a chart bar, an inner shape. The rest stays ink.
3.Add the rule (use nth-child(N) for a single accent, nth-child(n+M) to accent everything from index M onwards).
4.Add the icon to BiToneIconsSection's icons list in DesignSystem.tsx so the showcase covers it.
5.The default (no rule) is "all paths ink on active" — that is acceptable for icons with one strong silhouette.
Do not bi-tone icons that don't appear in the sidebar nav. Other locations (buttons, alerts, empty states) use the standard mono-color behaviour from §5.2.
6. Anti-patterns
Each item below has shipped at least once in this codebase or a sibling — that's why it's here.
•No raw hex anywhere. Not in style={{ color: '#…' }}, not in CSS, not in Tailwind arbitrary values like text-[#171717]. Always a token class.
•No pure black / pure white text. Use text-text-primary (#171717) and text-text-inverse.
•No new accent colors. Nothing beyond orange, ink, the greys and the muted danger red. accent-red / accent-lime resolve to ink; adding blue/teal/purple/green for "info" or "success" is forbidden; use the existing semantic tones (info is grey-family, success is dark text on light grey, warn is orange-family, danger is the muted red — the only red in the system).
•No orange-filled buttons. Orange is for state, focus rings, and highlight chips. Primary fill is cta-dark.
•No heavy shadows. shadow-card is the deepest non-drawer elevation. No shadow-2xl, no custom box-shadow with high alpha.
•No borders thicker than 1px. For more contrast, switch from border-border to border-border-strong.
•No tracking-tight. Use tracking-[-0.01em] / tracking-[-0.015em] exactly as shown in the typography showcase. tracking-tight is too aggressive.
•No font-bold. Use font-medium / font-semibold only. Bold reads as shouting in this system.
•No inline error rows. Failures surface through ErrorDialog (centered modal). See docs/ERRORS.md.
•No springs / bouncy motion. Ease-out / cubic-bezier only. See MotionList in DesignSystem.tsx.
•No style={{}} for colors, radii, shadows, motion. Token classes only. The one gradient in the system is the monochrome bloom — .mc-bloom / .mc-bloom-page in src/index.css (soft ink halos, no hue), used by GradientCallout, AuthBackground and the highlight StatTile. Reuse the classes; never re-declare radial-gradient values inline.
•No third-party UI libraries. See §3 anti-rule.
•No invented Pill tone, Switch tone, IconButton variant, etc. If a new variant feels needed, take the question to the owner before adding it.
•No fixed-width or `max-w- main canvas.** The content column right of the sidebar is a fluid 1fr` — see §2.8.
•No text-labelled buttons in table rows. Row actions are icon-only ghost IconButtons — see §3.2.
•No plain-text client/vehicle/case names. Every entity reference is a link to its detail route — see §3.3.
7. When this document is wrong
If you are working on a real screen and the rules above prevent a sensible solution, that is a signal — not permission to break the rules. Open a discussion (or, in agent mode, surface the conflict in your reply) before adding a new token, accent, or primitive. The system is intentionally narrow; every exception erodes consistency for the next agent.
The design contract in the harness

##Self-audit: the agent checks its own design

The contract alone isn’t enough. Coding agents drift: when they build a piece of frontend or a whole app, after a few screens they start working from memory instead of reading the Showcase. A “Create” button shows up in a top bar that doesn’t exist, the canvas turns sepia instead of white, a form makes up its own width. Every one of these actually happened. That’s why the harness has an audit skill.

Your Role
You are a design-system reviewer. You compare the implementation against the binding contract in docs/DESIGN_SYSTEM.md and the canonical patterns in src/pages/DesignSystem.tsx. You report concrete, citation-backed deviations. You do not fix them — your output is the audit report. Fixing is a separate pass (Step 5), so that the agent grading the work is never the agent that did it.
This skill exists because past builds shipped wrong chrome (CreateMenu in a topbar instead of the sidebar), wrong canvas color (bg-bg-base instead of bg-bg-surface), wrong form widths (max-w-[820px] instead of the templated max-w-[560px]), and freehand user footers that diverged from SidebarShellDemo. The build agent didn't catch these because it was synthesizing from memory instead of reading the showcase. Your job is to catch them after the fact.
Process
Step 1: Read the contract
1.docs/DESIGN_SYSTEM.md — §1 Agent contract, §2 App shell, canvas & locale, §3 Component selection rules, §5 Icon contract, §6 Anti-patterns. Internalize the rules — these are what you'll grade against.
2.src/pages/DesignSystem.tsx — walk it top to bottom. Pay special attention to:
–PageTemplatesSection — the three canonical layouts with exact widths, paddings, canvas color.
–SidebarShellDemo — the canonical app shell composition (CreateMenu in topSlot, user footer pattern, sidebar widths 260/76).
–CompositionSection — concrete reference patterns for detail page headers, KPI strips, form sections.
3.docs/COMPONENTS.md — the catalog of allowed primitives.
4.CLAUDE.md — the UI Build — Pre-flight section (the index that points at everything above).
5.src/index.css — the @theme block. This is the only source of color, radius, shadow, and font tokens.
Step 2: Pick the audit scope
Resolve the requested scope to a concrete file list:
•Feature ID → read knowledge/features/F{ID}-*.md, then git log --diff-filter=A --name-only or list the files mentioned in the spec to find what that feature added.
•Route path → trace from src/App.tsx to the page component(s).
•Glob → expand it.
•all / empty → every .tsx under src/pages/ (except DesignSystem.tsx) plus everything under src/components/ — the showcase-coverage check (§H) needs them in scope.
Print the resolved file list back to the user before reading them — short confirmation step.
Step 3: Run the audit
For each file in scope, check against this rubric. Cite specific line numbers when reporting a finding.
A. App shell & chrome (DESIGN_SYSTEM.md §2.1, §2.2, §2.4)
Verify against the rules in §2 of docs/DESIGN_SYSTEM.md and the canonical SidebarShellDemo in the showcase:
•One canonical AppShell exists; no freehand chrome in route files; no parallel shell implementations per portal.
•CreateMenu sits inside the sidebar topSlot; no topbar / header / floating create button.
•User footer matches the SidebarShellDemo pattern exactly (compare to the demo code).
•Sidebar widths match the demo (expanded / collapsed).
B. Page templates (DESIGN_SYSTEM.md §2.5)
For every route-level page, identify which PageTemplatesSection template it should follow (Dashboard / List + detail / Centered form). Flag pages that don't match a template, mix templates, or improvise a fourth shape without escalation.
C. Canvas color (DESIGN_SYSTEM.md §2.3)
•Main page canvas is bg-bg-surface (white). bg-bg-base (light grey) is allowed only inside drawer shells.
•No raw hex values anywhere. Grep for bg-\[ / text-\[ / border-\[ / style={{ with color values — these are anti-pattern per §6.
D. Tokens & anti-patterns (from §6)
•No font-bold (use font-medium / font-semibold).
•No tracking-tight (use tracking-[-0.01em] / tracking-[-0.015em]).
•No new accent colors — orange, ink, the greys and the muted danger red only (accent-red / accent-lime resolve to ink).
•No orange-filled buttons (orange is for state/focus, primary fill is cta-dark).
•No shadows heavier than shadow-card outside drawers.
•No borders thicker than 1px.
•No inline error rows — failures go through ErrorDialog.
•No third-party UI libraries (Headless UI, Radix, MUI, shadcn).
•No invented Pill tone, Switch tone, IconButton variant.
E. Component selection (from §3)
•Every JSX element should map to a primitive from docs/COMPONENTS.md or be composed from primitives. Flag freehand <div className="bg-bg-surface border ..."> that should be a Card or Surface.
•Every form field should be TextField or SegmentField, not a raw <input> styled by hand.
•Every status indicator should be Pill or StatePill, not a freehand pill-shaped div.
•Every section title should use PageHeader or SectionHeader, not a raw <h2 className="...">.
F. Icon contract (§5)
•Every icon's size / strokeWidth pair matches the table in DESIGN_SYSTEM.md §5.1 exactly — read the table, don't audit these from memory.
•No strokeWidth={1} (too thin) or strokeWidth={2.25}+ (too punchy).
•Icon color comes from a text-* class, never from style (§5.2).
G. Locale (DESIGN_SYSTEM.md §2.7)
Slovak, informal-direct register, no exclamation marks. Flag English copy, formal address ("Vy"/"Vám" outside of explicit legal copy), or exclamation marks.
H. Showcase coverage (only if scope includes src/components/)
•Every primitive in src/components/ has a Showpiece in src/pages/DesignSystem.tsx and a row in docs/COMPONENTS.md. Flag any new primitive missing from either.
Step 4: Emit the audit report
Produce a markdown report in this exact shape:
# Design audit — {scope} **Date**: {YYYY-MM-DD} **Scope**: {what was audited} **Files checked**: {N} ## Summary - {N} blocking issues - {N} should-fix issues - {N} nits ## Blocking (chrome / template / canvas — affects every page) ### B-1: {short title} - **File**: `src/components/AppShell.tsx:42` - **Rule**: DESIGN_SYSTEM.md §2.1 *App shell is one canonical component* - **Found**: {what the code does} - **Expected**: {what the showcase / docs say} - **Fix**: {one-sentence concrete fix} ### B-2: ... ## Should-fix (per-page deviations) ### S-1: {short title} ... ## Nits (cosmetic, low impact) ### N-1: ... ## Files audited - `src/components/AppShell.tsx` - `src/pages/portal/CustomerDashboard.tsx` - ...
Severity rules:
•Blocking = wrong chrome/shell, wrong canvas color, wrong page template, missing AppShell extraction, third-party UI library, raw hex, invented tokens. These break the system itself.
•Should-fix = wrong icon size, wrong heading element, freehand div where a primitive exists, form not centered when template requires centered, wrong sidebar width.
•Nit = inconsistent spacing within tolerance, slightly off copy register, missing showcase row for a primitive that already exists in production.
Step 5: Hand off, don't fix
This pass never edits code. The report is the deliverable — an auditor that fixes its own findings stops reporting them.
In an autonomous run the orchestrator takes it from here without asking: it opens a fresh pass, applies every Blocking and Should-fix finding, re-runs this audit on the same scope to confirm they are gone, and commits. Nits are applied only when that pass is already touching the file. The run does not continue to the next epic with Blocking findings open.
When a human is driving, the report stops here and they decide what gets applied.
Rules
•Cite line numbers for every finding. A finding without a citation is rejected.
•Reference the rule for every finding (which section of which doc). "Just feels off" is not a finding.
•No auto-fix in the audit pass. Even a one-character fix belongs to the follow-up pass (Step 5) — this pass reads, it never writes.
•Do NOT suggest creating new primitives to absorb deviations. The contract says "compose existing primitives" — if a deviation can't be expressed with existing primitives, that's an escalation, not a fix.
•Read the actual showcase for every check — don't audit from memory. The showcase is the source of truth; this file (CLAUDE.md, COMPONENTS.md) is the index.
•Keep the report under ~80 findings. If there are more, group by file and report top 10 per file with a note that more exist.
The design audit skill: what it reads, what rubric it grades by, and what shape its report takes. It doesn’t fix anything, it only reports.

At the end of every epic, the design audit runs as a separate subagent. It loads the design contract, the design system tokens, and the Showcase page, then goes through the screens that were built and compares them against those: shell and navigation, page templates, canvas color, tokens and anti-patterns, component selection, icons, localization, and whether every new component has a demo in the Showcase and a row in the catalog. Every finding has a file, a line, and the rule it broke. Without a citation, a finding doesn’t count. The audit doesn’t fix anything itself, it only reports, at three levels: blocking, should-fix, nit. The fix comes in the next step, and the epic isn’t done until it has zero blocking findings.

claude — ~/marcocar
~/marcocar (main) $ claude
> /sprint.design-audit F022
Scope: F022 — Case detail (Admin) · knowledge/product/features/F022-case-detail.md
Reading contract: docs/DESIGN_SYSTEM.md · src/pages/DesignSystem.tsx · docs/COMPONENTS.md · src/index.css
Resolved 4 files:
src/pages/admin/cases/CaseDetail.tsx
src/pages/admin/cases/CaseTimeline.tsx
src/components/CaseCloseModal.tsx
src/components/ServiceHistoryTable.tsx
# Design audit — F022 Case detail (Admin)
Date: 2026-07-14 · Files checked: 4
Summary: 1 blocking · 3 should-fix · 2 nits
## Blocking
B-1 Detail canvas uses sepia
src/pages/admin/cases/CaseDetail.tsx:38 rule DESIGN_SYSTEM.md §2.3
found <main className="bg-bg-base px-8 py-8">
expected bg-bg-surface — sepia is reserved for drawer shells
fix drop the wrapper; AppShell already owns canvas color and padding
## Should-fix
S-1 Freehand status pill
src/pages/admin/cases/CaseDetail.tsx:112 rule DESIGN_SYSTEM.md §3 · COMPONENTS.md
found <span className="rounded-full bg-lime-soft px-2 text-[11px]">Otvorená</span>
expected <StatePill state="open" />
S-2 Raw <h2> section title
src/pages/admin/cases/CaseTimeline.tsx:21 rule DESIGN_SYSTEM.md §3
found <h2 className="text-[15px] font-semibold">Časová os</h2>
expected <SectionHeader title="Časová os" />
S-3 Icon stroke off-contract
src/components/ServiceHistoryTable.tsx:64 rule DESIGN_SYSTEM.md §5.1
found <Wrench size={18} strokeWidth={2} /> inside a table row
expected size 16 · strokeWidth 1.75 for inline rows
## Nits
N-1 Exclamation mark in copy
src/components/CaseCloseModal.tsx:47 rule DESIGN_SYSTEM.md §2.7
found "Zákazka uzavretá!" ("Case closed!")
expected "Zákazka uzavretá" — informal-direct, no exclamation marks
N-2 Missing catalog row
src/components/ServiceHistoryTable.tsx rule DESIGN_SYSTEM.md §1.5
found Showpiece exists in the showcase, no row in docs/COMPONENTS.md
expected every primitive gets a showcase entry and a catalog row in the same commit
Report only — no files changed.
Say "fix blocking" to apply B-1 in a separate pass.
This is what the audit result for one epic looks like (an example). Every finding has a file, a line, and the rule it broke.

##Keeping the design system up to date

And the answer to who maintains the design system and the Showcase: the agent, because three rules force it to, and the audit checks them.

  • The 80 percent rule – For every piece of UI, it has to find the closest existing component, and if it’s at least an 80 percent fit, it uses it and tweaks it through props. It may not fork it.
  • Derivation recipes – When nothing fits, it follows the recipe for that shape: which sibling it inherits the API from, which tokens it may use, where to save the file. A new overlay can only be a drawer or a modal. There is no third kind.
  • “Show what you built” – Every new component gets, in the same commit, a demo in the Showcase and a row in the COMPONENTS.md catalog. The audit flags every one that’s missing them, so the catalog never falls out of sync with the code.
COMPONENTS.md 80% “show what you built” Showcase audit 80% fits → use it · doesn’t fit → derive by recipe · new → Showcase + COMPONENTS.md, or the audit flags it COMPONENTS.md 80% “show what you built” Showcase audit 80% fits → use it · doesn’t fit → derive by recipe new → Showcase + COMPONENTS.md, or the audit flags it

The anti-patterns section of the contract came straight out of these audits: every line in it is a bug from the calibration builds that the agent actually produced and that never happened again.

The result: when the agent built a layout or UI, it always knew how. Not once did it create UI at its own discretion. It always followed rules defined up front.

We have the design system. Now let’s look at how to define the harness so it generates a consistent frontend.

#The app’s frontend

##An app without a backend

With the design system done, we can move on to defining the frontend. Let’s leave the backend aside for now. We want to get the app into our hands as soon as possible, try it out, and iterate on what we see. At this point, a backend only adds complexity and time: time we’d have to spend defining it and the agent would spend building it.

But how do you get a working app into your hands without a backend? The trick is to simulate the backend on the frontend, which coding agents handle with ease. That’s why the first version of every app I build never has a backend: it’s frontend only, and the browser’s local storage serves as its database. On the web, that’s IndexedDB. Every feature that would otherwise live on the backend lives on the frontend for now. We don’t worry that it’ll have to be moved one day: once we sign off on the final shape of the app after X iterations, we add backend rules to the harness, throw the old app away, and start a new build with a backend – as easy as swapping a printer cartridge.

local DB client sees FE build · 4 hours backend prod DB tests · audit client doesn’t see not built yet local DB client sees FE build · 4 hours backend prod DB tests · audit client doesn’t see not built yet

##The harness, file by file

What does all this look like in the coding harness? Let’s go through it file by file. In the early phase of the project, when we were building only the frontend, the docs/ folder held these documents with predefined rules:

  • ARCHITECTURE.md – Describes what parts the app is made of and how they talk to each other: what a screen is, what logic is, and where data is stored. It also works as a living list of everything that gets built in the app over time, so both the agent and a human always see the current state.
  • BOUNDARIES.md – Sets the boundaries between the parts of the app – what may depend on what and what may not, so the code doesn’t get tangled. It also lists technologies and approaches we deliberately don’t use in the project, so the agent doesn’t even start adding them.
  • COMPONENTS.md – A catalog of all the interface building blocks (buttons, cards, tables, form fields) with a description of what each one is for. The agent is meant to assemble new screens from these pieces like building blocks, instead of inventing something new every time.
  • DECISIONS.md – A log of important decisions: what we decided and why. Thanks to it, nobody later questions or accidentally reverses a choice that was made for a reason.
  • DESIGN_SYSTEM.md – A binding manual for building the app’s look: how to approach designing a new screen, which visual rules are fixed, and what’s forbidden. It makes sure every screen looks as if the same designer drew it.
  • ROUTES.md – A map of all the app’s screens and the navigation between them, including which user role sees what. The agent uses it to know where a new screen fits and who gets access to it.
  • ERRORS.md – A single rule for how the app communicates errors to the user: always the same way, clearly, and without surprises. That way the user gets the same, predictable experience with every problem.

Take the list as inspiration, not a prescription. Every project needs different rules and different things to enforce on the agent.

These files are the alpha and omega of the rules the agent has to follow when it builds the app autonomously. I never put them in the agent’s main rules file, like CLAUDE.md, so as not to overload it. A better practice is to keep just an index of the files in CLAUDE.md or AGENTS.md: the agent reads the one that relates to the problem it’s solving right now, whether that’s architecture, new UI, or something else.

MarcoCar
Quick Reference
•Stack: React 19 + Vite + Tailwind CSS v4 + Dexie (IndexedDB)
•Knowledge base: knowledge/ (roadmap, feature specs, data model)
•Code docs: docs/ (design system, architecture, boundaries, components, decisions)
Critical Architecture Rules
•No backend — Everything runs locally. No server.
•Simulated auth — Login resolves a local UserAccount from IndexedDB; no server, no tokens, no password hashing. Roles gate the UI, not the data.
•IndexedDB via Dexie — Single source of truth. Include id, createdAt, updatedAt on every table.
•Local-first — App must work fully locally without backend.
Code Conventions
•TypeScript strict mode, no any
•Functional components + hooks only, no class components
•Custom hooks for all reusable logic (src/hooks/)
•Mobile-first, responsive design
•Tailwind v4 (CSS-first config, @import "tailwindcss") — Tailwind classes only, no CSS modules or CSS-in-JS
•Lucide React for icons
•Path alias: @/* maps to src/*
UI Build — Pre-flight
Before writing JSX for any route-level page, run this loop in order. Do not skip.
1.Read the contract. Open docs/DESIGN_SYSTEM.md — it is the binding contract, not background reading. The concrete rules (templates, canvas color, app shell, anti-patterns) live there.
2.Read the showcase. Open src/pages/DesignSystem.tsx and locate the section that matches what you're building (page templates, shell, composition examples). Mirror its structure verbatim — don't translate from memory.
3.Compose, don't fork. For each piece of the screen, pick the closest primitive from docs/COMPONENTS.md. If something fits within ~80%, pass props — don't write a new variant.
4.Never invent tokens. Colors, radii, shadows, motion durations come from src/index.css. If a value seems missing, you're trying to do something the system rejects — re-read the anti-patterns section of DESIGN_SYSTEM.md.
5.Audit before "done". Run /sprint.design-audit at the end of every build session. It is a hard gate, not optional polish.
When in doubt: docs/DESIGN_SYSTEM.md + src/pages/DesignSystem.tsx are the source of truth, not memory.
Commands
npm test && npm run lint
Project Structure
knowledge/ # Product knowledge base roadmap/ # The roadmap — epics, each split into features features/ # One spec per feature (F001, F002, ...) — goal, acceptance criteria, layout entities.md # Data model — every entity with its attributes, lifecycle and relations docs/ # Code documentation DESIGN_SYSTEM.md # The binding visual contract — read before any UI work ARCHITECTURE.md # System overview, data layer, state mgmt, hooks/services registry BOUNDARIES.md # Dependency flow + import rules per layer, hook & component rules COMPONENTS.md # Catalog of primitives, composites, and domain components DECISIONS.md # Technical decision log (Dxxx entries with rationale) ERRORS.md # Error handling pattern (centered ErrorDialog; toasts are success-only) ROUTES.md # Route table, navigation structure, deep-linking rules src/ # Source code components/ # Reusable UI components pages/ # Route-level page components hooks/ # Custom React hooks db/ # Dexie database schema and operations types/ # TypeScript interfaces services/ # Business logic data/ # Static data lib/ # Utilities test/ # Test setup
CLAUDE.md as the index, and behind it the harness’s docs/ folder, with longer files trimmed for this preview. The agent reads only the one that relates to its problem.

##A simulated backend and the seam under the hooks

Now back to simulating the backend. What matters is what it looks like inside. It’s not a pile of mock data in components. It’s a full-fledged database in the browser, written exactly the way it would look on the backend. The IndexedDB tables are a one-to-one copy of the data model from the harness: same entities, same names, every record has a UUID as its primary key, and relationships go through foreign keys. Even what can’t really happen in a browser, like sending an email or an SMS, gets simulated. Each send is logged as a record in the notifications table, and the link that would otherwise arrive by email shows up right on the screen. The case timeline, for example, works exactly the way it will for real, and the future backend has a seam ready to plug into (in reality it all gets rewritten, but who knows... it’s a test of good architecture design).

nothing above the zipper gets resewn seam components · pages · hooks same in both phases services · IndexedDB simulated backend in the browser phase 1 · local app Laravel services · Postgres same entities, same names production · zips in instead of IndexedDB nothing above the zipper gets resewn seam components · pages · hooks same in both phases services · IndexedDB simulated backend in the browser phase 1 · local app Laravel services · Postgres same entities, same names production · zips in instead of IndexedDB

To keep the simulation from leaking into the UI, the harness keeps the layers separate, and BOUNDARIES.md enforces it as an import rule: types, then the database, services with the business logic above it, hooks above them, and only on top of that, components and pages. A component or page may never import the database or a service. All data flows through hooks. This was written into the harness decisions from day one, for a single reason: when the backend arrives, the layer under the hooks changes, not the screens.

#Defining what gets built

##Roadmap and epics

We have the design system and the frontend rules. What’s left is the most essential part: what we’re building. That’s what the roadmap and its feature specs are for. The roadmap looks at the app from the top and defines the modules (and therefore epics) that need to be built to get a working app. Each module breaks down into several features, and each one has a description of how it works.

roadmap.md epics 64 specs feature spec goal interaction AC × 2–3 short, about the product, not the code roadmap.md epics 64 specs feature spec goal interaction AC × 2–3 short, about the product, not the code

The scope came out of the nine discovery workshops from the intro. The roadmap built from it has eight epics with 64 feature specs. Seven epics belonged to phase one. The eighth, fleets, came in phase two, but its scope was researched in those same workshops. No others were held.

##Specs: short and about the product

Each feature has its own .md document that describes how it should work. It’s not a classic specification like the ones you know from the early spec-driven frameworks. Mine has three parts: a Goal of three to five lines, two to three acceptance criteria, and a Layout & Design section on how the UI should be put together. For a purely backend feature, that third part is left out. People usually put more into feature specs: the data entity, its inputs and outputs, a technical specification, how to program the whole thing. I don’t do that. All the rules for how to implement something are written down once, in general terms, in the harness. There’s no need to repeat them for every feature. Today’s language models are smart enough to implement a feature on their own. Your job is to give the agent good rules at the coding harness level. Leave the rest to it.

That keeps the specs as simple as possible, so they’re easy to review, with no unnecessary detail. But one thing must never be missing: every feature and every screen states which user role it applies to. The agent won’t figure that out on its own, and in an app with four roles, “who can see and do what” is the most common source of bugs. That one sentence later grows into both the authorization tests and the user journeys.

Roadmap: MarcoCar
Status: Active · Data model: entities.md
Overview
React-based build of the admin portal and the customer portal. Eight epics, journey-ordered; the ones this excerpt touches: Foundation → Client Registration → Accident Intake → Case Lifecycle → Reporting → Fleet Portal. Role says who owns a feature — Admin, Customer, or Tech for platform work and background jobs with no user role.
<!-- Trimmed for the article — the earlier epics stand here in the real doc (and in yours). -->
5. Case Lifecycle
How a case moves from start to done. A case is born at the accident call (epic 4) and opened at intake, once the vehicle is on the bench. Ops manages it through the repair, then closes it with the final price and a short repair summary.
<!-- Trimmed for the article — the epic's remaining features stand here in the real doc (and in yours). -->
Features
IDFeatureRoleDescriptionStatus
F021All cases listAdminEvery case in one place. Filter by status and type.Built
F022Case detailAdminFull view of one case with its timeline; everything editable while open, including swapping the case type mid-flight.Built
F023Mark case doneAdminClosing dialog that captures the total repair price and a short repair summary, then flips the case to done.Built
Screens
RouteScreenPortal
/admin/casesAll cases listAdmin
/admin/cases/:idCase detail; the mark-done dialog opens from hereAdmin
<!-- Trimmed for the article — the remaining epics follow, one section each. -->
An epic from the roadmap and one of its specs, exactly as the agent reads them.

##The data model as a finished input

These files come with one more document, which belongs to the product, not to the code rules: the data model. It’s a single file, entities.md, that describes all the app’s entities along with their types. It’s over a thousand lines long, and it took shape during the workshops in the knowledge harness, long before the first build. Just like the roadmap and the specs, it went into the coding harness as a finished input, not something the agent would make up during the build.

roadmap · 64 specs coding harness brief: what to build data model entities.md · 1,000+ lines born in the workshops, long before the first build roadmap · 64 specs coding harness brief: what to build data model entities.md · 1,000+ lines born in the workshops, long before the first build

Every entity in it has the same structure. For example, a case:

  • Purpose – one sentence on what the entity is for in the system.
  • Attributes – a table of fields, each with its type, whether it’s required, which workshop or decision it comes from, and a note on what it means.
  • Lifecycle – the states and the transitions between them: pending intake → open → done.
  • Relationships – one case belongs to one vehicle and one client.

This is the very document the agent derived the IndexedDB tables from, and later the Postgres migrations for the backend.

Data Entities
Status: review · Sources: 9 client workshops + decisions + roadmap + specs
<!-- Trimmed for the article — the earlier entities stand here in the real doc (and in yours). -->
3.6 Case
Purpose: A reported accident and its repair job — the central object of the case lifecycle and of the monthly report.
Attributes
FieldTypeRequiredSourceNotes
case_iduuidyessystemPK
vehicle_iduuid (FK)yes—
client_iduuid (FK)yes—Resolved via vehicle
insurance_pathenum: pzp_victim / hp_comprehensive / otheryes (at intake)D011, D050Insurance alone decides it (D050)
case_statusenum: pending_intake / start / doneyesD015pending_intake = created at the accident call, waiting for intake · start = vehicle on the bench · done = client can pick up
accident_timetimestampyesD006, D050First-call time ≈ accident. Inline-editable while open
total_damage_eurdecimal(10,2)yes (at done)D014Entered by Ops when closing
case_open_time · case_done_time · softapp_job_id · repair_summary_comment · notes⋯
Lifecycle
pending_intake ──▶ start ──▶ done │ └─ created at the accident call (F018); waits for the intake, which opens it. case_type may flip mid-flight — must remain editable.
Email notifications fire when the case opens (→ start) and when it closes (→ done) — no email on creation.
Relationships
•1 Case → 1 Vehicle, 1 Client
•1 Case → 0..n ServiceRecord
The case in entities.md: purpose, attributes, lifecycle, relationships.

#Calibration: two builds until the prompt became a compiler

##The harness is the prompt

Once the design system, the code rules, and the roadmap with its specs are defined, the first test build can run. That’s the moment the work becomes tangible. If everything’s right, you’ll see a working app with all the UI elements, interactions, and features from the roadmap: fully clickable and finished-looking, built in a single autonomous run of the coding agent. The only prompt I wrote was “Build the whole app from the roadmap.” The harness is your prompt. That was just the trigger.

claude — ~/marcocar ~/marcocar (main) $ claude > Build the whole app from the roadmap ↵ ● Reading harness: CLAUDE.md · DESIGN_SYSTEM.md · roadmap.md ● Roadmap: 8 epics · 64 specs · data model ● Building the app, epic by epic ✓ 1 Foundation ✓ 2 Onboarding ⠋ 3 Client Registration · 4 Accident Intake · 5 Case Lifecycle · 6 Reporting · 7 Acquisition · 8 Fleet Portal ~4 h · go hit the gym meanwhile claude — ~/marcocar ~/marcocar (main) $ claude > Build the whole app from the roadmap ↵ ● Reading harness: CLAUDE.md · DESIGN_SYSTEM.md · roadmap.md ● Roadmap: 8 epics · 64 specs · data model ● Building the app, epic by epic ✓ 1 Foundation ✓ 2 Onboarding ⠋ 3 Client Registration · 4 Accident Intake · 5 Case Lifecycle · 6 Reporting · 7 Acquisition · 8 Fleet Portal ~4 h · go hit the gym meanwhile

##The first run is for calibration

The first run is usually a calibration run. It shows how well you’ve defined the design system, the code rules, and above all the roadmap with its specs. Click through the whole app, watch for deviations from what you intended, and trace every inconsistency back to the harness. If something’s off, chances are you’ll fix it by adding a rule, changing one, or defining what was left undefined. The same goes for the product specs: when something doesn’t work or doesn’t look the way it should, you’re usually one or two rule changes in the harness away from fixing it.

For me, it was small stuff: important components I’d forgotten to add to the catalog and the general rules (toasts, for example), a specific sidebar layout for different personas, or details of how form pages are put together.

build 1 build 2 build 3 fix the spec fix the harness ~95% identical from the same spec the first two builds weren’t about the app — they were about the harness build 1 build 2 build 3 fix the spec fix the harness ~95% identical from the same spec the first two builds weren’t about the app —they were about the harness

##Master harness: carry back what the agent figured out

Calibration runs have one more purpose. In the first build you see the whole app for real for the first time, and while building it, the agent also filled the gaps in your harness: it added components that weren’t in the catalog, wrote extra rules, made decisions. Some of them matter. So as the architect and orchestrator, after every calibration run you have to review what the agent added to the harness and the design system. Carry everything important or forgotten back into the starting state of the coding harness, the master version that every new build starts from. That way the harness isn’t calibrated only by your fixes, but also by what the agent figured out on its own.

master harness starting state of every build build + + calibration build the agent filled in what was missing review what goes back into the master toasts into the catalog sidebar per persona form composition one-off hack only what matters the rest goes out with the app what the agent figured out goes into the master calibration isn’t just fixing bugs — it’s also collecting what the agent figured out what the agent figured out goes into the master master harness starting state of every build build + + calibration build the agent filled in what was missing review what goes back into the master toasts into the catalog sidebar per persona form composition one-off hack only what matters the rest goes out with the app calibration isn’t just fixing bugs — it’s also collecting what the agent figured out

Then you throw the whole app away and start the run again. It costs you almost nothing, just a bit of time and some tokens from your subscription. And you’ll see for yourself: if your rules are well defined, the app will be 95 percent the same every time.

One of these builds, frontend only, took four hours. I recommend scheduling it as a background task and going to the gym or heading outside in the meantime.

##Iterating with the client

Congratulations, you now have a system where the input is a spec, the output is an app, and the app is disposable: throwing it away is cheaper than fixing it. Now you can go to the client, play with the app, ask for feedback, and iterate until you’re both happy. On this project I ran six frontend builds: the first two were for calibration, and the next four brought in major feedback from the client.

The real magic is that I didn’t implement the feedback by vibe coding right in the code, but at the level of the harness and the specs. I threw away the old app and generated a new one with the feedback already in it, without breaking any existing code.

In practice it looked like this: the client clicked through the app and found that the tables were missing column sorting. I added a rule to the design system: every table must have interactive column sorting. The next build already had it.

tables can’t be sorted 1 · client clicked through the app feedback DESIGN_SYSTEM.md + 2 · one line into the harness build 3 · the next build already had it tables can’t be sorted 1 · client clicked through the app feedback DESIGN_SYSTEM.md + 2 · one line into the harness build 3 · the next build already had it

Table sorting is a deliberately trivial example. Client feedback usually goes much further: a differently composed flow of screens, a different split of work between roles, a whole chunk of functionality nobody thought of in the workshops. But the path is always the same. The change goes into the harness and the specs, not the code, and the next build already contains it. And this is the biggest value of the whole system: big feedback costs as little as small feedback. In traditional development, the cost of a change is proportional to how much code already exists. Here it’s proportional to how many lines get added to the harness. The next chapter shows that this holds even for a change that breaks the data model.

#When a new module breaks the design: change without pain

One of the worst nightmares on a greenfield project: the product is almost done and you find out you have to rework part of it. It happened to me too. Not because the client wanted a new feature, but because the project had two phases, and while planning the first one I didn’t think through how the second phase’s module would fit into the solution. New scope arrived and broke the app’s current design.

##What happened

Phase two added the fleet module: a portal where a business customer manages their cars, appointments, and service history themselves. The vehicle, the key entity of the whole system, suddenly needed to belong to two worlds at once, internal operations and the fleet, and the phase-one data model hadn’t accounted for that. Some relationships had to be reworked. And the new module brought a whole new world of features into the admin, so the way the screens were split up stopped making sense. Exactly the state you know from ordinary products: a new feature gets added without refactoring the existing ones, and the interface starts to feel chaotic.

before internal operations vehicle fleets cracking new module bolted on two worlds pull at the vehicle, the UI is chaos after internal operations fleets + new module unaware of each other shared core vehicle · appointments · notifications · accounts the new module fits in one core, two modules, clean UI before internal operations vehicle fleets cracking new module bolted on two worlds pull at the vehicle, the UI is chaos after internal operations fleets + new module unaware of each other shared core vehicle · appointments notifications · accounts the new module fits in one core, two modules, clean UI

##A fix at the architect’s level, not in the code

In traditional development, this change would cost a week or two: someone has to map everything it touches and redo the data model, the migrations, the screens, and everything built on them. We did none of that in the code. The whole fix happened in the harness, at the highest level, where the architect decides:

  • Impact assessment. The agent went through the roadmap, all the specs, and the data model and mapped what the new module would break and which specs it would touch.
  • Change proposal. Together we designed a new core for the model: the vehicle as a shared entity, with two modules on top of it, internal operations and fleets, that are unaware of each other. The agent carried the change through the whole roadmap and all affected specs at once, so nothing we’d already carefully defined was left inconsistent.
  • UI cleanup. The admin navigation was regrouped into new sections so the new module would fit in, not be bolted on.

All in one day. The next day we threw the old app away, and the new build already had the fleet portal in it.

change fleet portal impact assessment vehicle has its owner baked in admin navigation no longer fits specs that break agent maps what breaks proposes roadmap.md + eighth epic · new admin navigation entities.md · data model +400 lines · core + 2 modules affected specs + notes · decision log +3 1 day in the harness · next day, a build with fleets change fleet portal impact assessment vehicle has its owner baked in admin navigation no longer fits specs that break agent maps what breaks proposes roadmap.md + eighth epic · new admin navigation entities.md · data model +400 lines · core + 2 modules affected specs + notes · decision log +3 1 day in the harness · next day, a build with fleets

This isn’t a story about a design flaw and its fix. It’s a feature of the system. A scope change that means weeks of pain anywhere else is something we can afford here, because we keep the app in its definition, not in the code. We’re architects: we make the change at the highest level and the code gets recompiled from it. It also helped that this happened in the frontend phase. In my process, the backend is always written at the very end. We had no code to be precious about. We had a definition we fixed, and an app we threw away and had rebuilt.

#Let’s add the backend and start the production run

##Switching the harness to a backend

Once the whole app design is done, every feature is in, the client is happy, and nothing’s missing, it’s time to switch the coding harness from “frontend only” mode to a full app with a backend and production deployment. This is a key checkpoint. It has to be well defined so that a single autonomous run can handle the whole app: the frontend, the backend, and the testing. For the client, nothing changes on the surface: the app looks the same as the one they clicked through. Only the waterline dropped, and everything else gets built below it.

local DB → migrations client sees same as before waterline in phase 1 backend · Laravel Postgres tests · audit deploy waterline dropped we build everything local DB → migrations client sees same as before waterline in phase 1 backend · Laravel Postgres tests · audit deploy waterline dropped we build everything

That means going through our coding harness in the docs/ folder and redefining the files that describe the architecture, the boundaries, and the other rules, from a local app to an app with a backend. I write backends in Laravel.

It’s the backend framework with the most complete ecosystem I know: one team builds both the core and official packages for everything a production app needs. Authentication and accounts (Fortify), queues and scheduled jobs (Horizon), emails and notifications, files, monitoring (Pulse, Nightwatch), testing (Pest), a connection to React with no API layer (Inertia), and hosting where the whole thing deploys with a single push (Laravel Cloud). For an agent, that’s ideal: every problem has one canonical, well-documented solution, so there’s no room to improvise. What’s left for me is defining the rules: how to serve the frontend, how to handle authentication and security, how to test the app.

##What survived from the local app

Nothing. Nothing was migrated. The frontend was written from scratch together with the backend, in one run, from the same specs. The only thing that survived from the local version was the design system: the components moved into the new project as a ready-made kit. The IndexedDB schema became Postgres migrations, and the hooks became controllers that send data to the pages via Inertia props. Every build, local and production, lived on its own git branch. Old ones were never edited. We just created a new one. The design phase wasn’t there to produce code we’d keep. Its job was to validate the app design and the data model with the client before we touched the backend and the final handover of the project. The code was a by-product, and it was disposable.

local app · phase 1 validated with the client product ✓ · data model ✓ design system · components the only thing that moves IndexedDB schema hooks pages code: a by-product specs design system · components Postgres migrations · new controllers · Inertia props · new pages · new production app · rebuilt from the specs local app · phase 1 validated with the client product ✓ · data model ✓ design system · components the only thing that moves IndexedDB schema hooks pages code: a by-product specs design system · components Postgres migrations · new controllers · Inertia props · new pages · new production app · rebuilt from the specs

##Two new documents

Compared to the frontend phase, the harness gained two new documents:

  • SERVICES.md – A binding “problem → package” map: every need the app has (authentication, queues, emails, PDFs, search, files) is assigned one official Laravel package or the framework itself. It’s the first-party packages rule from the stack section, written down so the agent doesn’t have to interpret it on its own.
  • USER_JOURNEYS.md – A catalog of every user journey across roles and portals, each with steps and an expected result. It’s a checklist for clicking through the whole app in the browser, role by role. I’ll show how it was used at the end of the chapter.

##Testing strategy and the verification gate

The testing strategy could easily have its own document in the harness. I put it as a general rule straight into CLAUDE.md, as the only exception to the rule that CLAUDE.md is just an index: it’s too important to make the agent go looking for it. The rule is short: every feature ships with tests in Pest, at minimum the happy path and authorization for every role. In an app with four roles, “who can see and do what” is the most common source of bugs, so this part is mandatory.

Then there’s the verification gate, six commands that have to pass before anything is declared done:

  • tests on SQLite, for speed,
  • the same tests again on Postgres, to match production,
  • backend code formatting,
  • a TypeScript type check,
  • lint,
  • the production build.

And one sentence that turned out to be the most important: the orchestrator runs the gate itself and never takes a subagent’s word that “tests passed.” I’ll explain who the orchestrator is in a moment.

##The orchestrator and two audits

The production run starts with the same prompt as every run before it: “Build the whole app from the roadmap.” The difference is how much the agent has to hold at once: rules for the frontend, backend, UI, UX, tests, and deployment, ten hours straight. For big runs like this, what’s worked for me is not launching a single agent, but an orchestrator that oversees the whole implementation process. For each epic in the roadmap, it launches implementation agents, waits for their reports, checks their work, launches testing and audit agents, and only when the epic is green and committed does it move on to the next one. It writes no code itself. It coordinates and validates.

✻ orchestrator Claude agent · 1M tokens of context launches reports agent · implementation code and tests for the epic agent · tests Pest · SQLite and Postgres agents · audit coding audit · design audit runs the gate itself, doesn’t trust the subagent verification gate · 6 commands commit → next epic epic by epic until the roadmap is done ✻ orchestrator Claude agent · 1M tokens of context launches reports agent · implementation code and tests for the epic agent · tests Pest · SQLite and Postgres agents · audit coding audit · design audit runs the gate itself, doesn’t trust the subagent verification gate · 6 commands commit → next epic epic by epic until the roadmap is done

There are two audits because we know agents don’t always follow the rules a hundred percent. You already know the design audit from the design system chapter: it checks whether the screens stick to the design contract, and whether every new component has a demo in the Showcase and a row in the catalog. The coding audit is new and does the same for the code: it goes through the epic’s frontend and backend and compares them against the rules in our coding harness. Both run as subagents after each epic is finished, the findings get fixed and committed, and only then does the run continue. That’s how you catch most of the cases where the rules aren’t followed. How do you create an audit skill like that? Either you write it by hand, or you let the agent analyze your coding harness and it writes the skill for you. The same goes for the testing strategy.

Or you don’t write anything at all and use existing skills from the framework’s creators. Laravel now publishes its own agent skills: an agent that cleans up freshly written backend code, plus skills for Laravel Cloud and Nightwatch, which cover exactly the deployment and monitoring from my stack. You install them in Claude Code as a plugin, and you’ve got a ready-made layer that guards the framework’s conventions.

##Ten hours without a single question

The production run took ten hours. It was one session running autonomously on my computer, with no interruptions and not a single question from the orchestrator: everything it needed to know was in the harness. The result: 30 migrations, 19 models, 33 controllers, 64 pages, 947 tests with 7,468 assertions, and four user roles and interfaces. A respectable scope for an app.

one session, ten hours questions: 0 · restarts: 0 · crashes: 0 0 h 5 h 10 h the only blips: a commit after each epic and out came 30 migrations 19 models 33 controllers 64 pages 947 tests · 7,468 assertions 4 roles and interfaces one session, ten hours questions: 0 · restarts: 0 · crashes: 0 0 h 5 h 10 h the only blips: a commit after each epic and out came 30 migrations 19 models 33 controllers 64 pages 947 tests · 7,468 assertions 4 roles and interfaces

I didn’t go in blind. We had six frontend runs behind us, so we knew the harness held. Before the live production run, there were two more backend test runs and one full-scope trial production run, just to see where the agents’ limits were at this scale and whether they could handle an autonomous task this long. It wasn’t an experiment but a calculated choice: from previous projects, I had plenty of experience with what agents can pull off. The test runs used Opus 4.8 first, then Fable 5. The difference in code quality wasn’t very noticeable. Both models followed the code rules well. All in all, the app was built ten times, and not once did anything crash or restart. No extra agent costs: the entire run was covered by my regular Anthropic subscription at €200 a month. The production run fit within the subscription’s five-hour usage-limit windows, if barely: when it rolled over from one window to the next, usage was at 99 percent, with the last minute ticking down.

the app was built ten times every run on its own git branch, old ones never edited €200 a month · regular subscription 6 × frontend 2 × backend test trial live nine rehearsals, one live run · 10 h the app was built ten times every run on its own git branch, old ones never edited €200 a month · regular subscription 6 × frontend 2 × backend test trial live nine rehearsals, one live run · 10 h

##Audits in action

The audits earned their keep. The coding audit caught, for example, a broken convention for writing backend controllers. The design audit found a page where the agent had invented its own header instead of using the existing PageHeader component, plus new components with no row in the catalog. The findings got fixed and committed, and the run moved on.

the implementation agent drifts from the rules course = rules in the harness audit audit audit controller convention own layout instead of the template component missing from the catalog after every epic the audit pulls it back, mid-run the implementation agent drifts from the rules course = rules in the harness audit audit audit controller convention own layout instead of the template component missing from the catalog after every epic the audit pulls it back, mid-run

##Tests and fifty clicked-through journeys

The agent wrote tests along with every feature, just as the rule in CLAUDE.md requires. The number of tests on its own says little. The number of assertions says how many things each test actually checks, and just under eight per test means the tests aren’t just checking that the page didn’t crash.

But tests can be written badly too, and an agent testing its own code tends to test what it wrote, not what it should have written. That’s why, after the production run finished, we started one more separate run: the agent got USER_JOURNEYS.md, opened a browser, and actually clicked through the app, role by role. That document wasn’t created after the run, but before it: the agent generated it from the roadmap and specs, I reviewed it, and from then on it mirrors everything you can do in the app, role by role. On top of that, the agent was instructed to act like a misbehaving user: try to break the app, enter nonsense, click where it shouldn’t. The same technique comes in handy later in the security review.

It logged in, walked through every flow, opened the downloaded attachments, and checked the sent emails. It fixed nothing, just logged every inconsistency. The result after fifty journeys: no blockers, no functional bugs in the app’s logic. It found two minor bugs, two inconsistencies in the demo data, and one place where the docs didn’t match the code. We fixed all of it afterward.

the builder inspects his own wall 947 tests · 7,468 assertions tests what it wrote and then final inspection: independent run, browser only 50 user journeys · every role log in, walk every flow attachments, sent emails fix nothing, just log it blockers 0 · logic bugs 0 minor bugs 2 · demo data 2 doc mismatch 1 an outsider who didn’t write the code the builder inspects his own wall 947 tests · 7,468 assertions tests what it wrote and then final inspection: independent run, browser only 50 user journeys · every role log in, walk every flow attachments, sent emails fix nothing, just log it blockers 0 · logic bugs 0 minor bugs 2 · demo data 2 doc mismatch 1 an outsider who didn’t write the code
User Journeys — browser smoke test plan
A complete catalog of user journeys across all portals. It serves as a repeatable checklist for the per-role browser smoke. Before a run: sail artisan migrate:fresh --seed (demo personas, password password for all of them: owner@ / ops@ / technik@ / fleet@ / zakaznik@marcocar.test).
ID convention: PUB (pre-auth/public) · OWN (admin as Owner) · OPS (Operations Manager) · TECH (Servisný technik) · FLT (B2B company customer, fleet) · B2C (individual customer) · X (cross-cutting).
OWN-12 — Fleets & customers (admin as Owner, owner@marcocar.test)
StepActionExpected
1Open /admin/fleetsList of portal customers (Customer): companies and individuals alike
2Switch the Firmy / Jednotlivci (companies / individuals, B2C) filterThe list filters by customer type
3Add a companyThe company is created in the invited state; a flash shows the activation link
4Open the company detailThe company's vehicles, service records and requests on one page
5Add a vehicle: first a free license plate (ŠPZ), then a taken oneThe free one goes through; the taken one shows a transfer checkbox and saves only after confirmation (confirm_transfer)
6Edit Termíny (due dates): date + reminderThe edit modal saves the date and toggles the reminder; a due date without a date stays „nezadané“ („not set“)
7Add a manual service recordThe record appears in the vehicle's História (history)
8Send the activationThe modal confirms it was sent and shows a one-time activation link, usable in PUB-5
<!-- Trimmed for the article — the other roles' journeys stand between these two. -->
B2C-3 — Booking a service (individual customer, zakaznik@marcocar.test)
StepActionExpected
1Log in and open /portalOverview of the single vehicle: the next due date and the „Objednať servis“ („Book a service“) CTA; with no vehicle, an EmptyState shows
2Open the vehicle detailDue dates with a status (green / orange / red) and the service history; the history shows only when the customer ↔ vehicle link is verified
3Book a service: type, two preferred dates, fault description, Mobilita (mobility)The request is sent; the platform stores no exact time, only preferences
4Open /portal/bookingsThe new request is in the „Čaká“ („Pending“) state
5Open its detail and cancel itrequested → cancelled; it disappears from the „Čaká“ filter and shows up under „Zrušené“ („Cancelled“)
6Try to open another customer's vehicle detail via a direct URLThey can't reach the detail — the customer ↔ vehicle link isn't verified, so the portal sends them back to its list
7Log out and open /portal/bookings via a direct linkRedirect to login; after logging in they land on their requests
<!-- Trimmed for the article — the real doc (and yours) has every journey, role by role. -->
Two of the user journeys the agent clicked through in the browser after the production run: steps and the expected result.

##Security review

In this process, security isn’t a separate discipline at the end, but the result of two decisions: which framework you pick and what you write into its rules. Laravel has most of the defenses built in, and the agent’s job is not to switch them off or get around them. Eloquent sends every query as a prepared statement with bound parameters, so SQL injection has no way through. Every write request needs a CSRF token, and Inertia sends it on its own. React escapes output, so XSS would need a deliberate bypass to get through. $fillable guards against mass assignment, a Form Request validates every input, and authorization lives in policies on the server – not in hidden buttons on the frontend. Fortify hashes passwords, rate-limits login attempts, and handles resets. Documents sit in private object storage, accessible only through an expiring signed URL. I didn’t invent any of it. I just wrote it into the harness as mandatory, so the agent couldn’t get around it, not even by accident.

The second layer was the agent as attacker. We ran the same misbehaving-user technique from the browser run once more, with the Fable 5 model and the reverse brief: break the app. It went through the known classes of attacks on a web app –

  • log in as a business customer and use guessed IDs and direct URLs to reach someone else’s case, vehicle, or fleet,
  • from a customer account, call admin routes and actions it has no right to,
  • slip extra fields into forms, SQL injection into filters and search, XSS payloads into text fields that get echoed back into the interface,
  • download someone else’s file from private storage via a tampered URL, reuse a session after logout, brute-force the login.

Result: nothing critical. Not a single attempt got where it shouldn’t have. It’s not luck, and it’s not the model’s doing: with four roles, “who can see and do what” is the most common source of bugs, which is why this process guards it with three layers at once – mandatory authorization tests for every feature, policies as a rule in the harness, and finally an agent that tries to get around it all.

reverse brief: break it agent as attacker someone else’s ID in the URL admin route extra form field SQL injection · XSS Laravel + harness rules policies CSRF · Form Request $fillable signed URL app still standing 0 breaches · nothing critical reverse brief: break it agent as attacker someone else’s ID in the URL admin route extra form field SQL injection · XSS policies CSRF · Form Request $fillable signed URL Laravel + harness rules app still standing 0 breaches · nothing critical

#Was the app done after the production run?

Functionally, yes. Every flow passed, every role got where it was supposed to. Handover-ready? Not yet: there were small things left that my trained eye would catch, not a test.

##Bugs that don’t deserve a rebuild

When I reviewed the code and tested the app, the production run turned out better than I’d expected. Of course, a few small bugs turned up. Looking back, they were mostly missing guidelines for the technical implementation, not logic errors in the code. For example, storing files on the local disk instead of external storage: after the next deployment, all the files would have been deleted.

The other bugs were in the same category. I’m picking them straight from the commit messages in git:

  • A “back” button in a form that accidentally submitted the form because it had no type set.
  • One login card with no inner padding, while all the others had it.
  • A modal that couldn’t be scrolled on short screens.
  • A percentage field that didn’t accept a decimal comma, only a period.
  • A text template that added a period after a value that already ended with one.
  • Ambiguous file naming.

Cosmetics, setup, small conventions. Not one of them was architectural, and that’s an important guideline for this whole approach. A bug deserves a rebuild with updated rules when it carries a fundamental problem in the architecture and, without a rule, would repeat on every screen that follows.

An example from our context: suppose the agent handled permissions on the frontend, hiding buttons and menu items by role, instead of on the server through policies. Every screen would inherit it, every one of them could be bypassed with a direct URL, and a fix in the code would mean going through all 64 pages. That bug calls for a rule and a rebuild. A bug that’s one-off and local doesn’t deserve a rebuild. It would be an expensive answer to a cheap problem.

bug repeats on every screen missing rule → harness rebuild with updated rules one-off and local fix in code a rebuild would be an expensive answer to a cheap problem bug repeats on every screen missing rule → harness rebuild with updated rules one-off and local fix in code a rebuild would be an expensive answer to a cheap problem

##You point at the bug, not the fix

So how did they get fixed? Not with vibe coding, where you dictate to the agent what to rewrite and where. The agent got the list of findings, just as the browser run and my own review had logged them, and worked through them autonomously: for each one it proposed a fix itself, implemented it by the harness rules, and ran the verification gate. I checked every fix before we saved it. You point at the bug, not the fix.

Personally, I think code today is better written than it was at the peak of my active engineering career, when I was building apps every day.

##The last human touch

The agents stuck to the design contract from the rules in the repo, and the UI felt solid to me. But as a retired UI/UX designer and a perfectionist, I spent roughly two more days polishing the UI and the flow, so that even the smallest detail and the smallest action in the app would feel pleasant and natural. For this part I have no carefully designed process, just good old vibe coding: Claude, do this like this, do that like that. It’s polishing small details, not vibe coding new features. A bit of the human touch the production version needed. It was the last human touch before handing over the project: a polish, not a rebuild.

vibe coding features “add a fleet portal” “export a case to PDF” “redo the vehicle intake flow” “add SMS notifications” we don’t do this: the change goes into the harness and specs, then a rebuild vibe coding details “back shouldn’t submit the form” “card: padding like the others” “make the modal scroll on mobile” “let percent accept a comma too” right in the code: “Claude, do this like this”, ~2 days before handover polishing small details, not new features vibe coding features “add a fleet portal” “export a case to PDF” “redo the vehicle intake flow” “add SMS notifications” we don’t do this: the change goes into the harness and specs, then a rebuild vibe coding details “back shouldn’t submit the form” “card: padding like the others” “make the modal scroll on mobile” “let percent accept a comma too” right in the code: “Claude, do this like this”, ~2 days before handover polishing small details, not new features

One autonomous run gets you an app that works, holds the design, and can be handed over after two days of polishing. That quality isn’t made in the run, but before it: in a harness with clear rules for how the app should look, what interaction pattern it has, and how it should be built. The last few percent are just a question of your eye for detail. And those get polished in the code, not in another run.

#What was left after the production run

Code. For now. There’s a point in my process where I switch over: the app becomes the single source of truth, and from then on we iterate on it, directly, if development is ongoing and the client keeps bringing new business requirements. The specs and the harness stay as documentation of why the app looks the way it does, but nothing gets compiled from them anymore.

##Six weeks instead of twelve

The whole execution phase of the project took six weeks. Before the era of coding agents, in my experience, the same scope took twelve weeks and needed a product manager, a designer, and two to three engineers. Today one person with a foot in product, design, and engineering can orchestrate it. And on top of that you get something we didn’t have before AI: the client has the app in their hands early in the solution design, feedback comes back in hours, and every build is the whole product – not a Figma prototype, not a piece of the app, not a stripped-down MVP. You don’t build code anymore. You have just two jobs: architect the right solution and orchestrate the agents that build it.

before: PM · designer · 2–3 engineers 12 weeks architecture 6 weeks · one person FE BE QA agents build whole product client · feedback in hours score = architecture · conductor = you · orchestra = agents · every build is the whole product before: PM · designer 2–3 engineers 12 weeks architecture 6 weeks · one person FE BE QA agents build whole product client · feedback in hours score = architecture · conductor = you orchestra = agents · every build is the whole product

#What was hard

##Managing the specs

The hardest part of the whole process is managing the specs for every feature in the roadmap. When you have, say, a hundred of them and you work with them all day, it’s genuinely exhausting, and you have to decide for yourself whether the trade-off is worth it. This is where you need the full power of agents: let them help you manage the specs. And keep them lean and simple so they don’t grind you down later: today’s agents aren’t dumb, and even from minimal information they’ll understand what you need. Always keep the key information at the harness level, so you put as few duplicate instructions into the specs as possible.

#12#48#97#64 a hundred specs, all day… managing a hundred specs is exhausting: let agents manage them and keep them lean #12#48#97#64 a hundred specs, all day… managing a hundred specs is exhausting: let agents manage them and keep them lean

##A complex UX flow as a prototype

The second hard thing is describing a more complex UX flow in a spec, say a ten-step process the user has to go through. Describing it well enough that the agent replicates it the way you picture it is hard. Especially when you recompile the app again and again in the iterative phase. What worked for me was to take the design system and iterate my way (yes, by vibe coding) to a standalone prototype with a UI that does exactly that tricky flow. I saved it in the harness under docs/prototypes, and the spec for that feature contains only a reference: the agent looks at it and implements it in the production code cleanly and by the production rules. That way you get the best of both worlds, vibe coding and spec as code.

flow prototype vibe-coded from the design system 12345 109876 ten steps words can’t describe saved in the harness: docs/prototypes/ reference feature spec Goal Acceptance Layout & Design → see docs/prototypes/ agent production code clean, by the harness rules the prototype shows how, the spec points to it, the agent builds it by the production rules flow prototype vibe-coded from the design system 12345 109876 ten steps words can’t describe saved in the harness: docs/prototypes/ reference feature spec Goal Acceptance Layout & Design → see docs/prototypes/ agent production code clean, by the harness rules the prototype shows how, the spec points to it, the agent builds it by the production rules

#One level up

##The prompt, wait, check loop

I’m noticing that a lot of engineers are losing the joy they used to get from their work because of AI agents. Programming was their everyday work, and AI is taking the joy out of it. An engineer’s typical day used to go like this. You get up in the morning, make coffee, open your IDE, and dive into a problem for eight hours: you break it into parts, build a solution, and by evening it works. You created something, solved something, and that was your reward. Today the same person sits all day in a prompt, wait, check, prompt loop. They make one small decision after another, the agent does the work in between, and by evening they’re exhausted from deciding, with no sense of having built anything.

before: eight hours deep in a problem morning evening evening: it works, you built something today: in the loop all day prompt wait check evening: exhausted, built nothing before: eight hours deep in a problem morning evening evening: it works, you built something today: in the loop all day prompt wait check evening: exhausted, built nothing

##The work has changed, and so has the reward

I see it differently: you have to go one level up. We don’t build the app at the code level anymore. We build the harness: the knowledge that steers the agent, the design system, the tests, and the audit, so the agent builds the app from scratch in one autonomous run. It’s different work, and judgment moves higher: what technologies to build with, what the UX and UI should look like, how the system should behave when something is missing. And you can play and iterate with it just like with code. The reward just lives somewhere else. Not a feature that works, but a system that works. And what comes out of it is three times bigger than what we’ve built so far.

The question is no longer how to build a piece of code well. The question is how far to raise your ambitions. To keep that reward from disappearing, the scope of what we can do has to grow. Once, we hunted game, then sowed fields, then wrote code that automated one thing. Today we can design a whole system that grows and changes with what the world around it wants from it. That’s in our hands now: to design it and orchestrate it.

ambition → ↑ reward hunt game the catch sow a field the harvest write a piece of code one thing works build an app the product works build a harness the system works: app ×10 a system that grows adapts on its own you today the reward stays when the scope grows: not a feature that works, but a system that works ambition → ↑ reward hunt game the catch sow a field the harvest write a piece of code one thing works build an app the product works build a harness the system works: app ×10 a system that grows adapts on its own you today the reward stays when the scope grows: not a feature that works, but a system that works

#Where this is heading

##The harness is disposable too

This article is more about a mental model than an exact process. It shows that it can be done this way, not that it has to be done exactly this way. Don’t copy my harness documents. Take inspiration and try what fits you: every person and every project has their own ways of working and is at their own stage. And the harness you build today won’t be relevant in six months. It’s nonsense to believe it will.

What’s true for the app is true for the harness: it’s disposable. Only one thing has to stay: the willingness to throw away the old mental model when the world changes, and build a new one. Today, AI itself will build the harness for you. What it won’t build for you is flexible thinking and the courage to experiment.

start here the world changes new model, new tool mental model bends, doesn’t break new harness AI writes it for you old harness throw away spec build app throw away months · when the model, tool, or ambition changes hours · ten times per project recompile, don’t refactor works one level up too: the harness is to the mental model what the app is to the harness start here the world changes new model, new tool mental model bends, doesn’t break new harness AI writes it for you old harness throw away spec build app throw away months · when the model, tool, or ambition changes hours · ten times per project recompile, don’t refactor works one level up too: the harness is to the mental model what the app is to the harness

##The product builder in the field

Here’s where I’m heading. Building greenfield projects fast, where a design flaw isn’t a disaster but a cheap fix. Where a feature added mid-flight isn’t a curse, it’s a feature. Where I, as a product builder, spend more time in the field talking with the customer and the users than sitting at the drawing board and in front of coding screens. And where the vision of a solution can be realized a lot sooner and more cheaply than was possible before. Not a stripped-down MVP, but a finished product.

at the drawing board and coding screens alone, with code time moved here in the field, with the customer and users you customer and users “we’d need this differently” product builder in the field: a design flaw is a cheap fix, a feature added mid-flight is a feature at the drawing board and coding screens alone, with code time moved here in the field, with the customer and users you customer and users “we’d need this differently” product builder in the field: a design flaw is a cheap fix, a feature added mid-flight is a feature

##Next experiment: fixed data, replaceable app

And one thing I want to try next time. Is the final production code really the source of truth, or is it still the product spec? I’m toying with the idea of separating out the production database as a static element and leaving the app as a dynamic element that can be recompiled again and again, even through further big iterations. Whether you add a feature or remove one, you always recompile the app. Picked up a huge number of users and the app can’t keep up? You recompile the backend with a different architecture, or in a different language or environment. The production data your customers and your company generated is fixed. Everything above it is replaceable. I’ll save that for future experiments, maybe for a future article.

spec app data spec v1 spec v2 · + new feature spec v3 · 10× the users recompile recompile recompile app v1 thrown away app v2 · + feature thrown away app v3 · different language different architecture when it can’t keep up … production data · fixed artifact · only grows, never thrown away code isn’t the source of truth: spec on top, data below, a swappable app in between spec app data spec v1 spec v2 · + new feature spec v3 · 10× the users recompile recompile recompile app v1 thrown away app v2 + feature thrown away app v3 different language different architecture when it can’t keep up … production data · fixed artifact only grows, never thrown away code isn’t the source of truth: spec on top, data below, a swappable app in between

And that’s just one idea. As I write this, a huge number of problems and opportunities come to mind that could be reinvented this way. I’m sure that while reading, you came up with at least ten things for another ten articles. I’m not afraid that AI will leave us with nothing to do. Quite the opposite: there’ll be a lot more work. It’ll just be different.

If you’re building something similar, like your own harness, a spec as the source of truth, or a disposable app, get in touch. I’d be glad to compare notes.

Share this articleLinkedInX
Peter Papp
Peter Papp

AI-era evangelist. I show what one person can do today: I build generative AI solutions and share what I learn along the way. Get inspired, take what’s useful, and go build. Working on something similar? Message me.

Be that one person

Building something with AI, or wondering where it could help your business? I’d love to talk it through.

Fastest way to reach me: LinkedInMessage me

or

© 2026 Peter Papp · Based in Slovakia ·
Built with ☕ and Claude Code
LinkedInXSlovenčina