# 02 — Architectural Mastery

**Goal:** make foundational decisions deliberately — before they become expensive — and bootstrap scalable apps (React / Next.js 15 App Router / strict TS) with a repeatable blueprint.

> **The core principle:** Architecture is the set of decisions that are **expensive to reverse**. Everything else is implementation detail. The job is to identify those decisions, defer the reversible ones, and make the irreversible ones with a written rationale.

---

## Part 1 — Mental models (think in these before any decision)

### 1.1 The reversal-cost lattice
Every decision has a reversal cost. Classify before choosing:

| Decision | Reversal cost | Example |
|---|---|---|
| **Cheap to reverse** | Hours–days | Component naming, folder layout, hook API, styling approach *within* a framework |
| **Moderate** | Weeks | State library, data-fetching layer, routing structure |
| **Expensive** | Months+ | Framework choice, monolith vs micro-frontends, server/client rendering paradigm, API contract shape, DB schema exposed to the client |

**Rule:** never spend architectural ceremony on cheap-to-reverse decisions (YAGNI lives here); spend it on expensive ones. If a decision is expensive, write the ADR (01).

### 1.2 Conway's law, exploited
> "Organizations design systems that mirror their communication structure."

The architecture will match your team topology whether you like it or not. **Design for the team you have and the growth you expect:** if 3 frontend engineers own one app, a modular monolith with clear boundaries beats micro-frontends. If 4 teams must ship independently to one product surface, boundaries (module federation / MFEs / package isolation) become a team problem, not a tech preference. Never adopt micro-frontends for a team that fits one repo — you are buying coordination costs you don't have.

### 1.3 Architecture is a trade-off lattice, not a ladder
Every "best practice" is a trade against something. Use the lattice:

| You want | You pay | Deciding question |
|---|---|---|
| Server rendering (SEO, fast first paint) | Server cost, cache complexity, hosting constraints | Does the page need SEO/linkability or is it an authenticated app? |
| Client rendering (cheap, simple) | Slow first paint, no SEO, weaker perf on low-end devices | Is it an internal tool / app behind login? |
| Micro-frontends (team independence) | Bundle duplication, shared-state friction, design-system governance | Do 3+ teams ship independently to the same surface? |
| Big design system investment | Upfront months, governance overhead | Are there ≥3 products/teams building UI? |
| Edge/global deployment | Debugging complexity, data locality | Is the audience global or regional? |
| Full type-safety end to end (OpenAPI/Zod codegen) | Codegen pipeline setup, schema governance | Do backend + frontend move fast enough that contract drift hurts? (Almost always yes) |

### 1.4 The three-box mental model (before writing any code)
1. **Domain box:** What is this system *fundamentally*? (A dashboard, a checkout, a collaboration tool…) — this decides state shape and data model.
2. **Delivery box:** How does code reach users? (CDN, edge, regions, offline/PWA needs) — decides rendering and caching.
3. **Evolution box:** What will change in 18 months? (New team members? Mobile app consuming the same API? Multi-tenancy?) — decides boundaries and contracts.
Decide each box independently, then compose. Most bad architecture is mixing the three (e.g., choosing a data-fetching library because of a delivery concern).

### 1.5 The decision matrix (ask before every foundational call)
Score options 1–5 per row, weight the rows that matter to *your* context:

| Criterion | Weight (1–5) | Option A | Option B |
|---|---|---|---|
| Team familiarity & hiring | | | |
| Reversal cost if wrong | | | |
| Performance ceiling (your metrics) | | | |
| Ecosystem maturity & longevity | | | |
| Operational complexity (hosting, CI) | | | |
| Fit with backend/org direction | | | |
| **Weighted total** | | | |

**Rule:** if the matrix is within 10% and you have no data, pick the **more boring** option — boring scales in hiring, debugging, and longevity. Never pick an architecture to impress an interviewer; pick one that survives a team change.

---

## Part 2 — The bootstrap blueprint (step-by-step)

### Step 0 — Contract the problem (30–60 min, written)
- One-paragraph purpose; the **primary user journey**; the **success metric** (LCP/INP budget, conversion, DAU — pick one).
- Write the **ADR-0**: scope, non-goals, team size, 18-month evolution guess (box 3). Everything downstream references this.

### Step 1 — Set up the monorepo (pnpm workspaces or Turborepo)
- One repo, clear package boundaries: `apps/web` (Next.js), `packages/ui` (design system), `packages/config` (shared tsconfig/eslint), `packages/api-client` (generated client), `packages/utils`.
- Root `tsconfig` strictness baseline inherited everywhere (see Step 4).
- **AI-assisted:** generate the workspace skeleton with Claude Code in a sandbox — prompt: *"Scaffold a pnpm monorepo with Turborepo containing a Next.js 15 app router app and three packages (ui, config, api-client) with strict TS."* Review every generated file before committing — the AI is a fast typist, not an architect.

### Step 2 — Next.js 15 App Router foundations
- **Route groups** for layout scopes: `(marketing)` (static, SSG), `(app)` (authenticated, dynamic), `(auth)` (login flows) — each with its own layout and rendering strategy.
- **Server-first by default:** data fetching in Server Components; `'use client'` only at interaction boundaries. If a component doesn't handle events or hold state, it should be a server component.
- **Streaming:** use `loading.tsx` + Suspense boundaries per route segment — never block the shell.
- **Caching semantics — know them precisely:** `fetch` caching defaults, `revalidateTag`/`revalidatePath` for ISR, and the `staleTimes` client router cache in Next 15. Document the cache map per data type in ADR-1.
- **Server Actions** for mutations inside the app; keep them thin (validate → call service → revalidate) and never trust client input.

### Step 3 — Define the data contract FIRST (contract-first)
- Backend publishes an **OpenAPI spec**; frontend **generates the typed client** (openapi-typescript + a thin fetch wrapper, or `orval`). No handwritten API types — they drift.
- Runtime validation at the **boundary only** (zod on server action inputs / route handlers); everywhere else types are trusted because they were generated.
- Error model agreed with backend: stable error codes (`NOT_FOUND`, `RATE_LIMITED`…), one `ApiError` shape, typed per endpoint. QA and backend sign off on this document (this is 01 influence in action).

### Step 4 — Strict TypeScript, zero-exception
```jsonc
// packages/config/tsconfig.base.json
{
  "compilerOptions": {
    "strict": true,
    "noUncheckedIndexedAccess": true,
    "exactOptionalPropertyTypes": true,
    "noImplicitOverride": true,
    "noFallthroughCasesInSwitch": true,
    "noUnusedLocals": true,
    "noUnusedParameters": true,
    "verbatimModuleSyntax": true,
    "erasableSyntaxOnly": true,
    "moduleDetection": "force",
    "skipLibCheck": false
  }
}
```
- `noUncheckedIndexedAccess` alone removes an entire class of "silly bugs" (04). Add `any`-ban in ESLint (`@typescript-eslint/no-explicit-any: error`), with a 0-tolerance policy: a suppressed rule requires a written comment with a ticket.

### Step 5 — Data layer (server state discipline)
- **One server-state library** (TanStack Query) or server components + router cache — not both doing the same job in the same area.
- Query keys as a typed factory (`queryKeys.user.profile(id)`) — stale keys are the #1 source of "data not refreshing" bugs.
- Mutations: optimistic updates with snapshot rollback (see the pattern in the Scale & UX section) or invalidation — decide per mutation and write it down.

### Step 6 — Design system seed (don't boil the ocean)
- Token-first (colors/spacing/type as CSS variables), one Button/Input/Modal/Dialog accessible by default, Storybook, and a **governance note** (when to add vs compose). Never build 40 components before the product has 3 screens — YAGNI is a *cost* rule, not a purity rule.

### Step 7 — Observability from day one (not week 12)
- Error tracking with source maps; RUM for Web Vitals (INP/LCP/CLS); one analytics event spec (`event, properties` typed). You cannot architect what you cannot measure — the first "performance work" is installing the meter.

### Step 8 — CI as the gate
- Lint → typecheck → unit tests → build → e2e (smoke on critical journeys) → **bundle-size check** (size-limit) → deploy preview. The pipeline is the enforcer of everything in 04.

### Step 9 — Rollout & release discipline
- Feature flags from the start (even a local one) so any release is reversible; **canary deploys**; a written rollback drill. Reversibility is an architectural property — cheap reversals let you take bigger swings.

---

## Part 3 — AI-assisted architecture (make it a force multiplier)

- **Use Claude Code/Copilot for research and skeletons, never for the decision.** Prompt for: *"Compare [options] for [context]; list trade-offs with sources; produce a decision matrix"* — then you apply judgment.
- **Generate the ADR draft from your decision**, then edit — the discipline is the artifact, not the prose.
- **Ask for the counter-argument:** after choosing, prompt *"argue against this decision as a skeptical principal engineer"* — free red teaming before you commit.
- **Sandbox spikes:** generate a spike implementation in a throwaway branch to validate an unknown (e.g., "does module federation work with our auth flow?") — timebox to a day, then delete or formalize.

---

## ✅ The architecture checklist (before any foundational commit)
- [ ] Reversal cost classified; expensive decisions have an ADR
- [ ] Three-box model answered (domain / delivery / evolution) in writing
- [ ] Decision matrix scored for any framework/library choice ≥ moderate reversal cost
- [ ] Team topology considered (Conway's law) — architecture fits the org
- [ ] Data contract-first: OpenAPI → generated typed client; error model agreed
- [ ] Strict TS baseline on; `any` policy enforced in CI
- [ ] Rendering strategy per route group documented (SSG/ISR/SSR/CSR)
- [ ] Observability (errors + RUM) installed before feature work
- [ ] Feature flags + rollback drill exist
- [ ] The boring option won unless data said otherwise

→ Next: [03 — Flawless Execution](03-flawless-execution.md) · Back: [01 — Visibility & Influence](01-visibility-and-influence.md)
