The studio
The systems your team is trying to build.
We already run them.
Most AI is a demo that passed. Ours answers a public phone line, gates a supply chain, and runs a back office in production, under one law: AI reads, reasons, and proposes; code enforces what it may change, spend, or send; every consequence leaves a receipt. We build the same class of system for you, on infrastructure you own, in weeks.
Or call our own production voice agent and try to break it: +1 (617) 766-0577.
What you can buy
Fixed scope, fixed price, published below. Half to schedule, half on delivery: if it is not deployed and working, you do not pay the back half.
We run it for you
Not a tool you learn. A desk that runs. Audits of what you already have, and the watch that keeps it correct.
Voice Agent Production Readiness Audit
$5,500 · 5 business daysYour voice agent survives the demo. We audit whether it survives real callers: every tool call and webhook on the path, auth and secrets at the edge, booking races and calendar truth, CRM capture and deduplication, transcript persistence, failure visibility, and prompt-injection posture on public caller input. Ranked findings, critical fixes shipped or precisely scoped, and a regression checklist your team keeps.
For: agencies and teams running Vapi, Retell, or ElevenLabs agents in front of customers. Ours answers +1 (617) 766-0577. Call it. Try to break it. The full audit page →
Book this audit →AI Codebase Debt Scan
$495 · 2 business daysFive to ten confirmed findings from your repository in two business days: duplicated logic, CWE-class security issues, dead code, and architectural drift, each traceable to a file and line. Anything we could not confirm is left out, and the report says how many findings were dropped rather than counting them.
For: teams that want the number before the audit. The full ledger and the paydown quote are the next rung. The debt page →
Book a scan →AI Codebase Debt Audit
$1,500 · credited against the fixAI wrote a serious share of your codebase and nobody is accountable for it. We scan the repo and hand you a debt ledger with a dollar sign on it: duplication, CWE-class vulnerabilities, dead code, and architectural drift, every finding traceable to a file and line. Ranked by what it costs you monthly, with a fixed price for the paydown and the full fee credited if we do the work.
For: teams that shipped fast with Cursor, Copilot, or Claude Code and can feel the maintenance bill arriving. Industry data says duplication is up and refactoring has collapsed; your repo has its own numbers, and we put receipts on them. The first repo we audited was our own. The debt page →
Book this audit →Rescue Diagnosis
$1,500 · credited against the fixYour AI project stalled, hallucinates, or got abandoned mid-build. We diagnose where it actually broke: the model, the tool wiring, the data, or the design. You get a written diagnosis, a ranked fix plan, and a real number for the repair, with the full fee credited if we do the fix.
For: teams whose AI vendor ghosted them, and internal builds that never made it to production. How a rescue runs →
Book a diagnosis →Production Operations Retainer
from $3,500 · a monthFor a system we built or audited. Watchdogs on the seams that fail quietly: the public door, the destination tables, the webhook registry, the config a vendor can rewrite overnight. A receipt on every consequential action. A nightly reconcile that pages a person when a transcript or a row goes missing. A monthly report with numbers you can audit, and one live fix window a month under the same gates our own builds pass. Reporting, billing reconciliation, and model-cost routing ride inside it.
For: the agent that still has to be right in month four. Month to month, cancel any month; the report is the receipt. The operations page →
Book a scoping call →We build it for you
The same law, installed on your premises: the model reads and proposes, the database decides what it may do, the audit trail keeps the receipts.
AI Document Pipeline
$14,500 · 2-3 weeksEmail/portal intake → AI classification & extraction → deterministic rule engine (AI never auto-approves) → human-review queue → writes to your system of record. AV scanning, idempotent processing, audit trail, 30-day warranty.
For: invoices, COIs, claims, intake packets, contracts, or any document flow a human is drowning in.
Book this pipeline →Production Agent Build
$8,500 · 2 weeksOne agent workflow, end to end: your integrations (CRM, email, Slack, DB), guardrails and injection-hardening, observability, dead-letter handling, deployed on your infra or ours.
For: triage, enrichment, follow-up, reporting, or the workflow your team does manually every day.
Book this build →AI Enablement Sprint
$3,500 · 1 weekYour dev team, upgraded: Claude Code setup with persistent memory and search, MCP tooling, CI guardrails, two live working sessions, and a written playbook.
For: teams who bought the tools but aren’t getting the velocity.
Book this sprint →Agent Platform Build
from $25,000 · 4-6 weeksThe thing your team has been trying to stand up internally: a control-plane skeleton, your first two production agents, persistent memory layer, governance rails (database-enforced rules, audit trails, human override paths), and the playbook to add agent three yourselves.
For: companies past the chatbot phase. We already run this pattern, shown in the flagships on the proof page.
Book a scoping call →Every engagement: fixed scope, fixed price, 50% to schedule, 50% on delivery. If it isn’t deployed and working, you don’t pay the back half. The retainer is the one exception and says so: month to month, cancel any month, the report is the receipt.
What people ask us for
Fourteen asks a mid-market buyer makes, in the buyer’s own words, and the offer that answers each one. Fourteen asks, five shapes: check it, run it, build it, teach it, remember it. If you run a staffing firm, the whole desk is on one page.
- An AI audit and ROI blueprint before we spend a dollar
- The free working session, then the $1,500 diagnostic credited against the build
- One context layer instead of twenty disconnected tools
- The memory layer inside the Agent Platform Build
- Client onboarding that runs without a human touching it
- Production Agent Build, or the Document Pipeline when intake arrives as paper
- Reporting pulled from CRM, ads, and finance with no manual input
- Production Agent Build, kept honest under the operations retainer
- Tier-one agents for service, scheduling, and intake
- The front desk we run ourselves, built for you as an Agent Build; audited if you already have one
- Replace overlapping subscriptions with apps on our own data
- Agent Platform Build
- Deal ingestion, enrichment, and staleness alerts
- Production Agent Build
- Proposals, contracts, and memos drafted from structured data
- Production Agent Build
- Scattered spreadsheets and inboxes into one queryable system
- AI Document Pipeline
- A knowledge base trained on our own SOPs
- The memory layer, delivered in the Sprint or the Platform Build
- Who is over or under capacity, in real time
- Production Agent Build
- Catch scope creep and invoice errors before they leak revenue
- AI Document Pipeline, reconciled monthly under the retainer
- Get the team to use what we bought
- AI Enablement Sprint
- Cut inference spend without cutting quality
- Model routing, reported monthly under the retainer
How an engagement starts
- Working session · 30 minutes, free. Your system and what it would take to ship. You leave with direction whether or not you hire us.
- Diagnostic · optional, $1,500 flat. A deeper two-hour workflow map and a written build plan you keep, credited in full against your first build if we proceed. For teams who want the full picture before committing to a pilot.
- Fixed-scope pilot. One priced offer from the ladder above, deployed in weeks. If it is not deployed and working, you do not pay the back half.
- Operate or hand over. We keep the system correct under the Production Operations Retainer, or your team takes the playbook and owns it. You own the code, infra, prompts, and docs either way.
Every engagement is founder-led. Your session is with Cap.
How we ship
Every build runs the same rail: scope, spec, adversarial review, test-gated build, live verification. Every phase gates on a failing test written first. A green run that skipped the work is treated as red. Adversarial reviewers attack each change before it merges, and verification gates refuse unproven claims. Agents investigate, propose, build, and challenge. A person authorizes anything irreversible and owns the consequences. The rail, and why it is fast, in one essay: The Bottleneck Was Never the Builder.
SCOPE / SPEC / ADVERSARIAL REVIEW / TEST-GATED BUILD / VERIFY LIVE
Why our systems survive production
- AI does perception. Code does verdicts. Humans take the edge cases. Our systems never let a model auto-approve anything load-bearing. Extraction is probabilistic, decisions are deterministic, and anything inconclusive routes to a person.
- Speed with receipts. The portfolio above is dated. Open any marked number to inspect its method, scope, and verified date. How the receipts are kept honest: Verify From Where the Caller Stands.
- You own everything. Code, infra, prompts, docs. No platform lock-in, no rented black box.
- The agent is the continuity. Models get swapped. The history, the commitments, and the receipts stay, and the authority to act is enforced outside the model. That is the category we build in, and the reason month four looks like month one.
About
Gnosis Labs is an independent human-agent systems studio. Cap Dawes owns delivery and every consequential commitment; the persistent agents that work beside him investigate, design, build, challenge, and carry the operating history from one session to the next. We run production fleets and cognitive runtimes of our own design, and we build the same class of system for clients. Thirteen years buying, grading, and delivering technology; AI research and technical operations since 2021; the production agent portfolio since February 2026.
Questions buyers ask
What does Gnosis Labs build?
Gnosis Labs builds production AI agents, document automation pipelines, and agent platforms for companies. The studio also runs its own multi-agent systems in production, with persistent memory and database-enforced governance, and ships its own consumer AI products. Every system is multi-tenant, guarded, monitored, and deployed, not a demo.
How can Gnosis Labs deliver in weeks what agencies quote in months?
Every build runs through an in-house agent fleet: specialized builder agents in isolated workspaces, adversarial reviewers that attack each change before it merges, and verification gates that refuse unproven claims. Agents investigate, propose, build, and challenge; a person authorizes anything irreversible and owns the consequences. That removes calendar time without cutting corners.
How much does an AI agent or document pipeline cost?
Pricing is fixed and published. An AI Codebase Debt Scan is $495, a Rescue Diagnosis is $1,500 (credited against the fix), an AI Codebase Debt Audit is $1,500 (also credited against the fix), an AI Enablement Sprint is $3,500, a Voice Agent Production Readiness Audit is $5,500, a Production Agent Build is $8,500, an AI Document Pipeline is $14,500, and an Agent Platform Build starts at $25,000. The Production Operations Retainer is from $3,500 a month, for a system we built or audited. Half is due to schedule and half on delivery on the builds; the retainer is month to month.
How long does an AI build take?
Most engagements ship in one to six weeks. A sprint is one week, an agent build is two weeks, a document pipeline is two to three weeks, and a full agent platform is four to six weeks.
Do I own the code and infrastructure?
Yes. You own the code, infrastructure, prompts, and documentation. There is no platform lock-in and no rented black box.
How do you keep AI systems reliable in production?
AI handles perception, deterministic code makes the decisions, and humans handle the edge cases. Models never auto-approve anything load-bearing, and inconclusive cases route to a person. Systems are AV-scanned, idempotent, and fully audited.
What has Gnosis Labs already shipped?
Live systems include a production front desk that answers our own phone line and books real meetings; OurPool, a real-time World Cup pool platform; Side, a consumer quiz and affiliate engine; Glassbridge, a stateful AI game-master engine; and Marrow, the semantic memory engine under the studio’s own agent fleet. Behind them: DawnForged, an event-sourced RPG engine, and the client systems described on this page.
What is the Production Operations Retainer?
A monthly engagement, from $3,500, for a system we built or audited. We put watchdogs on the seams that fail quietly, keep a receipt on every consequential action, run a nightly reconcile that pages a person when a transcript or a row goes missing, send a monthly report with numbers you can audit, and hold one live fix window a month under the same gates our own builds pass. Reporting, billing reconciliation, and model-cost routing ride inside it. Month to month; the report is the receipt.
What is an AI Codebase Debt Scan?
A $495 scan of a repository that AI helped write, delivered in two business days: five to ten confirmed findings, each traceable to a file and line, covering duplicated logic, CWE-class security issues, dead code, and architectural drift. Anything we could not confirm is left out, and the report says how many findings were dropped. It is the entry rung to the full Debt Audit.
What is a Rescue Diagnosis?
A fixed-price $1,500 diagnosis of an AI build that stalled, hallucinates, or was abandoned by another vendor. We find where it actually broke: the model, the tool wiring, the data, or the design. You get a written diagnosis, a ranked fix plan, and a real cost for the repair, and the full $1,500 is credited against the fix if we do the work. Broken projects are a specialty, not a chore.
What is an AI Codebase Debt Audit?
A fixed-price $1,500 audit of a codebase that AI helped write. Nothing has to be visibly broken: the point is what is accruing quietly. We measure duplication, CWE-class vulnerabilities, dead code, and architectural drift, then hand you a debt ledger where every finding is traceable to a file and line and the total carries a dollar figure. You get a ranked paydown plan with a fixed price on the fix, and the full $1,500 is credited if we do the work. Independent 2026 research finds AI-assisted codebases duplicating more and refactoring less; your repo has its own numbers, and this puts receipts on them.
What is a Voice Agent Production Readiness Audit?
A fixed-price, five-business-day audit of a deployed voice agent: the full call path from telephony through tool calls, webhooks, calendar booking, and CRM capture. We attack the failure paths a demo never hits: wrong HTTP methods, silently dropped parameters, stale tool configurations, booking races, orphaned transcripts, and prompt injection through caller speech. Our own receptionist answers +1 (617) 766-0577 on exactly this discipline. $5,500, half to schedule. Delivery and acceptance are explicit: the ranked findings report, the agreed critical fixes shipped or precisely scoped with evidence, and a regression checklist your team keeps.