The studio

The systems your team is trying to build.
We already run them.

Most AI is a demo that passed. Ours answers a public phone line, gates a supply chain, and runs a back office in production, under one law: AI reads, reasons, and proposes; code enforces what it may change, spend, or send; every consequence leaves a receipt. We build the same class of system for you, on infrastructure you own, in weeks.

Or call our own production voice agent and try to break it: +1 (617) 766-0577.

What you can buy

Fixed scope, fixed price, published below. Half to schedule, half on delivery: if it is not deployed and working, you do not pay the back half.

We run it for you

Not a tool you learn. A desk that runs. Audits of what you already have, and the watch that keeps it correct.

Voice Agent Production Readiness Audit

$5,500 · 5 business days

Your voice agent survives the demo. We audit whether it survives real callers: every tool call and webhook on the path, auth and secrets at the edge, booking races and calendar truth, CRM capture and deduplication, transcript persistence, failure visibility, and prompt-injection posture on public caller input. Ranked findings, critical fixes shipped or precisely scoped, and a regression checklist your team keeps.

For: agencies and teams running Vapi, Retell, or ElevenLabs agents in front of customers. Ours answers +1 (617) 766-0577. Call it. Try to break it. The full audit page →

Book this audit →

AI Codebase Debt Scan

$495 · 2 business days

Five to ten confirmed findings from your repository in two business days: duplicated logic, CWE-class security issues, dead code, and architectural drift, each traceable to a file and line. Anything we could not confirm is left out, and the report says how many findings were dropped rather than counting them.

For: teams that want the number before the audit. The full ledger and the paydown quote are the next rung. The debt page →

Book a scan →

AI Codebase Debt Audit

$1,500 · credited against the fix

AI wrote a serious share of your codebase and nobody is accountable for it. We scan the repo and hand you a debt ledger with a dollar sign on it: duplication, CWE-class vulnerabilities, dead code, and architectural drift, every finding traceable to a file and line. Ranked by what it costs you monthly, with a fixed price for the paydown and the full fee credited if we do the work.

For: teams that shipped fast with Cursor, Copilot, or Claude Code and can feel the maintenance bill arriving. Industry data says duplication is up and refactoring has collapsed; your repo has its own numbers, and we put receipts on them. The first repo we audited was our own. The debt page →

Book this audit →

Rescue Diagnosis

$1,500 · credited against the fix

Your AI project stalled, hallucinates, or got abandoned mid-build. We diagnose where it actually broke: the model, the tool wiring, the data, or the design. You get a written diagnosis, a ranked fix plan, and a real number for the repair, with the full fee credited if we do the fix.

For: teams whose AI vendor ghosted them, and internal builds that never made it to production. How a rescue runs →

Book a diagnosis →

Production Operations Retainer

from $3,500 · a month

For a system we built or audited. Watchdogs on the seams that fail quietly: the public door, the destination tables, the webhook registry, the config a vendor can rewrite overnight. A receipt on every consequential action. A nightly reconcile that pages a person when a transcript or a row goes missing. A monthly report with numbers you can audit, and one live fix window a month under the same gates our own builds pass. Reporting, billing reconciliation, and model-cost routing ride inside it.

For: the agent that still has to be right in month four. Month to month, cancel any month; the report is the receipt. The operations page →

Book a scoping call →

We build it for you

The same law, installed on your premises: the model reads and proposes, the database decides what it may do, the audit trail keeps the receipts.

AI Document Pipeline

$14,500 · 2-3 weeks

Email/portal intake → AI classification & extraction → deterministic rule engine (AI never auto-approves) → human-review queue → writes to your system of record. AV scanning, idempotent processing, audit trail, 30-day warranty.

For: invoices, COIs, claims, intake packets, contracts, or any document flow a human is drowning in.

Book this pipeline →

Production Agent Build

$8,500 · 2 weeks

One agent workflow, end to end: your integrations (CRM, email, Slack, DB), guardrails and injection-hardening, observability, dead-letter handling, deployed on your infra or ours.

For: triage, enrichment, follow-up, reporting, or the workflow your team does manually every day.

Book this build →

AI Enablement Sprint

$3,500 · 1 week

Your dev team, upgraded: Claude Code setup with persistent memory and search, MCP tooling, CI guardrails, two live working sessions, and a written playbook.

For: teams who bought the tools but aren’t getting the velocity.

Book this sprint →

Agent Platform Build

from $25,000 · 4-6 weeks

The thing your team has been trying to stand up internally: a control-plane skeleton, your first two production agents, persistent memory layer, governance rails (database-enforced rules, audit trails, human override paths), and the playbook to add agent three yourselves.

For: companies past the chatbot phase. We already run this pattern, shown in the flagships on the proof page.

Book a scoping call →

Every engagement: fixed scope, fixed price, 50% to schedule, 50% on delivery. If it isn’t deployed and working, you don’t pay the back half. The retainer is the one exception and says so: month to month, cancel any month, the report is the receipt.

What people ask us for

Fourteen asks a mid-market buyer makes, in the buyer’s own words, and the offer that answers each one. Fourteen asks, five shapes: check it, run it, build it, teach it, remember it. If you run a staffing firm, the whole desk is on one page.

An AI audit and ROI blueprint before we spend a dollar
The free working session, then the $1,500 diagnostic credited against the build
One context layer instead of twenty disconnected tools
The memory layer inside the Agent Platform Build
Client onboarding that runs without a human touching it
Production Agent Build, or the Document Pipeline when intake arrives as paper
Reporting pulled from CRM, ads, and finance with no manual input
Production Agent Build, kept honest under the operations retainer
Tier-one agents for service, scheduling, and intake
The front desk we run ourselves, built for you as an Agent Build; audited if you already have one
Replace overlapping subscriptions with apps on our own data
Agent Platform Build
Deal ingestion, enrichment, and staleness alerts
Production Agent Build
Proposals, contracts, and memos drafted from structured data
Production Agent Build
Scattered spreadsheets and inboxes into one queryable system
AI Document Pipeline
A knowledge base trained on our own SOPs
The memory layer, delivered in the Sprint or the Platform Build
Who is over or under capacity, in real time
Production Agent Build
Catch scope creep and invoice errors before they leak revenue
AI Document Pipeline, reconciled monthly under the retainer
Get the team to use what we bought
AI Enablement Sprint
Cut inference spend without cutting quality
Model routing, reported monthly under the retainer

How an engagement starts

  1. Working session · 30 minutes, free. Your system and what it would take to ship. You leave with direction whether or not you hire us.
  2. Diagnostic · optional, $1,500 flat. A deeper two-hour workflow map and a written build plan you keep, credited in full against your first build if we proceed. For teams who want the full picture before committing to a pilot.
  3. Fixed-scope pilot. One priced offer from the ladder above, deployed in weeks. If it is not deployed and working, you do not pay the back half.
  4. Operate or hand over. We keep the system correct under the Production Operations Retainer, or your team takes the playbook and owns it. You own the code, infra, prompts, and docs either way.

Every engagement is founder-led. Your session is with Cap.

How we ship

Every build runs the same rail: scope, spec, adversarial review, test-gated build, live verification. Every phase gates on a failing test written first. A green run that skipped the work is treated as red. Adversarial reviewers attack each change before it merges, and verification gates refuse unproven claims. Agents investigate, propose, build, and challenge. A person authorizes anything irreversible and owns the consequences. The rail, and why it is fast, in one essay: The Bottleneck Was Never the Builder.

SCOPE / SPEC / ADVERSARIAL REVIEW / TEST-GATED BUILD / VERIFY LIVE

Why our systems survive production

About

Gnosis Labs is an independent human-agent systems studio. Cap Dawes owns delivery and every consequential commitment; the persistent agents that work beside him investigate, design, build, challenge, and carry the operating history from one session to the next. We run production fleets and cognitive runtimes of our own design, and we build the same class of system for clients. Thirteen years buying, grading, and delivering technology; AI research and technical operations since 2021; the production agent portfolio since February 2026.

Questions buyers ask

What does Gnosis Labs build?

Gnosis Labs builds production AI agents, document automation pipelines, and agent platforms for companies. The studio also runs its own multi-agent systems in production, with persistent memory and database-enforced governance, and ships its own consumer AI products. Every system is multi-tenant, guarded, monitored, and deployed, not a demo.

How can Gnosis Labs deliver in weeks what agencies quote in months?

Every build runs through an in-house agent fleet: specialized builder agents in isolated workspaces, adversarial reviewers that attack each change before it merges, and verification gates that refuse unproven claims. Agents investigate, propose, build, and challenge; a person authorizes anything irreversible and owns the consequences. That removes calendar time without cutting corners.

How much does an AI agent or document pipeline cost?

Pricing is fixed and published. An AI Codebase Debt Scan is $495, a Rescue Diagnosis is $1,500 (credited against the fix), an AI Codebase Debt Audit is $1,500 (also credited against the fix), an AI Enablement Sprint is $3,500, a Voice Agent Production Readiness Audit is $5,500, a Production Agent Build is $8,500, an AI Document Pipeline is $14,500, and an Agent Platform Build starts at $25,000. The Production Operations Retainer is from $3,500 a month, for a system we built or audited. Half is due to schedule and half on delivery on the builds; the retainer is month to month.

How long does an AI build take?

Most engagements ship in one to six weeks. A sprint is one week, an agent build is two weeks, a document pipeline is two to three weeks, and a full agent platform is four to six weeks.

Do I own the code and infrastructure?

Yes. You own the code, infrastructure, prompts, and documentation. There is no platform lock-in and no rented black box.

How do you keep AI systems reliable in production?

AI handles perception, deterministic code makes the decisions, and humans handle the edge cases. Models never auto-approve anything load-bearing, and inconclusive cases route to a person. Systems are AV-scanned, idempotent, and fully audited.

What has Gnosis Labs already shipped?

Live systems include a production front desk that answers our own phone line and books real meetings; OurPool, a real-time World Cup pool platform; Side, a consumer quiz and affiliate engine; Glassbridge, a stateful AI game-master engine; and Marrow, the semantic memory engine under the studio’s own agent fleet. Behind them: DawnForged, an event-sourced RPG engine, and the client systems described on this page.

What is the Production Operations Retainer?

A monthly engagement, from $3,500, for a system we built or audited. We put watchdogs on the seams that fail quietly, keep a receipt on every consequential action, run a nightly reconcile that pages a person when a transcript or a row goes missing, send a monthly report with numbers you can audit, and hold one live fix window a month under the same gates our own builds pass. Reporting, billing reconciliation, and model-cost routing ride inside it. Month to month; the report is the receipt.

What is an AI Codebase Debt Scan?

A $495 scan of a repository that AI helped write, delivered in two business days: five to ten confirmed findings, each traceable to a file and line, covering duplicated logic, CWE-class security issues, dead code, and architectural drift. Anything we could not confirm is left out, and the report says how many findings were dropped. It is the entry rung to the full Debt Audit.

What is a Rescue Diagnosis?

A fixed-price $1,500 diagnosis of an AI build that stalled, hallucinates, or was abandoned by another vendor. We find where it actually broke: the model, the tool wiring, the data, or the design. You get a written diagnosis, a ranked fix plan, and a real cost for the repair, and the full $1,500 is credited against the fix if we do the work. Broken projects are a specialty, not a chore.

What is an AI Codebase Debt Audit?

A fixed-price $1,500 audit of a codebase that AI helped write. Nothing has to be visibly broken: the point is what is accruing quietly. We measure duplication, CWE-class vulnerabilities, dead code, and architectural drift, then hand you a debt ledger where every finding is traceable to a file and line and the total carries a dollar figure. You get a ranked paydown plan with a fixed price on the fix, and the full $1,500 is credited if we do the work. Independent 2026 research finds AI-assisted codebases duplicating more and refactoring less; your repo has its own numbers, and this puts receipts on them.

What is a Voice Agent Production Readiness Audit?

A fixed-price, five-business-day audit of a deployed voice agent: the full call path from telephony through tool calls, webhooks, calendar booking, and CRM capture. We attack the failure paths a demo never hits: wrong HTTP methods, silently dropped parameters, stale tool configurations, booking races, orphaned transcripts, and prompt injection through caller speech. Our own receptionist answers +1 (617) 766-0577 on exactly this discipline. $5,500, half to schedule. Delivery and acceptance are explicit: the ranked findings report, the agreed critical fixes shipped or precisely scoped with evidence, and a regression checklist your team keeps.

Receipt

Verified
How it's measured
What it doesn't claim

No receipt, no number. Every figure on this site resolves here.