Articles liés à AI Agent Architecture: The Verifiable Core: Guardrails,...

AI Agent Architecture: The Verifiable Core: Guardrails, Invariants, and Deterministic Design — a Core That Disposes of Every Proposal, Evals with a ... ... (The AI Agent Architecture Series) - Couverture souple

Livre 6 sur 7: The AI Agent Architecture Series

MOBILUCK - CODE247.AI, VU TRI CONG

 
9798173887870: AI Agent Architecture: The Verifiable Core: Guardrails, Invariants, and Deterministic Design — a Core That Disposes of Every Proposal, Evals with a ... ... (The AI Agent Architecture Series)

Synopsis

The model proposes. The core disposes. The pack is the proof.

Atlas, Meridian Supply Co.'s customer-operations agent, passed its exam and its week by Book 5. Then a pharmaceutical wholesaler's audit letter asked a question none of the tests could answer: what can this system never do, and what is the evidence? This book is the answer — a verifiable core built one failure at a time, each from an incident, each measured against an adversary, each proved with a test that turns green.

Every chapter is a lab on the companion repository — one dependency, fully offline with a scripted mock, every listing printed from a verified line range, every command paired with its expected output. Five moves per chapter: Run the demo, Read the listing, Break it with the chapter's planted failure, Fix it with the design move, Prove it with a test that turns green. You will:

  • Write the impact map — twenty-one failure modes scored on severity, likelihood, and detectability — and compute which ones the suite witnesses: fourteen before the book, twenty-one after
  • Put a deterministic core between the model's proposal and the world, ten rules judged against the decision, and run a rogue model through it: fifty-two refusals, zero sentences out
  • Turn the top of the map into invariants over the world — exactly one mail, no refund above the limit without approval, no other account's data — generated into eighty-six cases and green under three models
  • Build a wall at each of a run's three doors — input, actions, output — and read the table that says what each wall alone lets through
  • Make evals infrastructure: thresholds in a file the review diffs, a gate with an exit code, and a model judge calibrated at 81 percent and kept out of the gate
  • Draw the threat model per entry point, find the two walls that were missing, and write six attacks that succeed with the defences down and fail with them up
  • Trace every run as a span tree with a signature, record the model's and the world's replies, replay to the same hash, and name the first divergent span after a change
  • Model the refund protocol in forty lines and check one invariant over every reachable state: a counterexample in thirty-three states without the key, none in 14,872 with it

Every chapter carries a research lineage (FMEA and STPA, Anderson's reference monitor, Floyd–Hoare and QuickCheck, Saltzer–Schroeder and the OWASP list, Shewhart and the judged judge, STRIDE and the confused deputy, Dapper and record-and-replay, Nygard and fail-safe design, model checking from Clarke–Emerson to Amazon's TLA+, safety cases and supply-chain attestations), a five-item failure catalog, an applied deep-dive, and exercises in three tiers. Ten figures, 300 tests, fifty-three decision records, one running system that grew from Book 5's team without changing a node of it.

Who it's for: engineers and architects who have an agent that works and need to show — to a security review, an auditor, a regulator, or themselves — what it can never do. Assumes Books 2–5 or the equivalent; basic Python; no framework, no GPU.

The AI Agent Architecture Series is the architect's track: one architectural layer per book, on one running system the reader refactors and grows by hand. This is Book 6, The Verifiable Core — trust, designed and proved.

Les informations fournies dans la section « Synopsis » peuvent faire référence à une autre édition de ce titre.