Skip to content
back to blog

3 minEngineering

Building SoloSteady: applied AI where the numbers have to be right

SoloSteady turns twelve months of bank deposits into an income summary for self-employed people applying for a bank-statement mortgage. I build and run it solo under my LLC. I set out to learn applied AI and agentic systems in production. The biggest lesson was where an AI model belongs in a money document, and where it does not.

Live: SoloSteady

Case study: SoloSteady

Built and run solo under my LLC, Culmination Digital. Source is private.

SoloSteady makes one document. A self-employed person connects their bank, the app reads twelve months of real deposits, and it produces an income summary they can hand to a lender for a bank-statement mortgage, the kind of loan underwritten on deposits instead of tax returns. The buyer downloads it and delivers it themselves. The product never sends anything to a lender. It ships in English and Spanish.

What I wanted to learn

  • Applied AI. Putting a model inside a real product, not a demo.
  • Agentic systems. Orchestration, retrieval (RAG), and the Model Context Protocol.
  • Agentic systems in production. Cost caps, timeouts, failure handling, and proof that the model behaves.
  • AI-assisted, spec-driven development. Writing the decisions and rules first, then building with a coding agent inside them.

The lesson: keep the model away from the numbers

The domain set the boundary. An income figure a lender will read cannot come from a model. It has to be reproducible, explainable line by line, and identical every time it is computed.

So the analysis is deterministic. Deposits are classified by rules, grouped by month, and totaled by code. Every number in the document comes from that code, and a stored analysis built by an older version of the rules is recomputed rather than served.

Claude does one job: it writes the plain-language paragraph that explains the result, in both languages, from figures the code already computed. That one job is fenced in on every side:

  • The call is forced into a schema and validated with Zod before anything uses it.
  • An output guard rejects forecasting, approval, or cause language on every request, and any warning the model writes is replaced with fixed, translated wording.
  • A spend cap records the cost of every call, and each call has a hard timeout.
  • An eval suite with a judge model runs in CI before promotion to production, and a weekly scheduled canary runs it again against the live system.

The orchestration, retrieval, and MCP practice went into my open-source work, where an agent's mistake costs nothing: Forge fans a stack trace out to four specialist agents in parallel, Kev-O is a retrieval-grounded assistant with its own public evals, and fieldops-mcp is an MCP server. SoloSteady is where I learned what production asks of all of that: a model is one component with a contract, and everything around it has to hold when it misbehaves.

How it's built

Next.js 16 on Vercel, TypeScript strict, Postgres, Plaid for the bank connection, Stripe for checkout, the Anthropic SDK, Resend for email, Sentry for errors, Upstash for rate limiting, and pdf-lib for the document. Bank access tokens live in an encrypted vault, every webhook signature is verified, row-level security is on the tables, and admin access requires a second factor.

GitHub Actions runs lint, typecheck, a copy gate, unit tests and deterministic evals, integration tests against a local database, end-to-end tests with accessibility checks, and the model evals. The copy gate exists because of the domain: it blocks any wording that would make the product sound like it reports on a person to a lender, and it fails the build if an English string has no Spanish twin.

Spec-driven, with an AI pair

SoloSteady has 112 recorded architecture decisions, a set of rule files that tell a coding agent how the analysis pipeline, the API routes, the financial data, and the copy must work, and nearly 400 unit test files. I write the decision first, the agent builds inside it, and the tests and evals decide whether it is done. When a decision turns out to be wrong, it gets superseded in writing rather than quietly changed, so the reasoning survives.

What I took away

Applied AI in production is mostly everything around the model: the contract, the guard, the eval, the cap, and the decision record that explains why the model is allowed to do exactly one thing.