pot/labs

// code meets commerce

We build and run agentic AI for enterprise. Then we try to break it.

Pot Labs is a South African firm that designs, builds and operates agentic systems — agents that do a named job inside your estate, workflows that close themselves, and the adversarial harness that proves either one actually works.

The commercial unit is an artefact with an exit criterion, not a person on a desk for a quarter. An engagement that finishes, finishes.

POT LABS (PTY) LTD· JOHANNESBURG· AGENTIC AI · ASSURANCE

/01 · practice areas

Four practice areas. Three we have built.

Each is a named engagement with a delivery shape, a stated dependency on you, and an exit criterion written as a command rather than an adjective.

/01

AI Agents

custom-built · deployed in your environment

An agent that does a named job inside your system, on your infrastructure, under your credentials. You do not buy "an AI capability" — you buy a running system plus the three artefacts that make it maintainable without us.

  • A constitution: versioned invariants the system may not violate
  • The harness that proves them — adversarial suites wired into CI
  • An operating model telling the next agent how work here is organised
/02

Managed Agents

operated by us · on a clock · against a live system

A standing graph of scheduled runs, each with a named output, plus the evidence trail it produces. Not a support retainer and not a seat count.

  • Nightly adversarial suites; findings tracked with an owner, not a dashboard
  • Weekly spec-drift detection — contracts win, the code is wrong
  • An evidence artefact that regenerates on command and exits non-zero on failure
/03

Agentic Workflows

headless · intake to close · no person in the happy path

A commercial process that completes without a human on the success path — and where that property is machine-asserted rather than observed. Most of the design work is deciding what not to automate.

  • A close that asserts it completed unattended, and fails loudly when it did not
  • Every decline carries a registry code and a trace
  • The refusal list: what this workflow will never do without a person, and why
/04

Adversarial Assurance proposed

gates that block · checkers that did not write the code

Sold standalone, against a system we did not build. A single command that decides whether work is done, and a dated, commit-stamped report your risk function can read.

  • Distinct lenses — correctness, abuse, reproduction, contract, external standard
  • A skipped test counts as a failure, enforced in CI rather than by inspection
  • Every report carries a register titled "what this does not yet prove"

This practice area is proposed on internal evidence. It has never been sold, and we will not pretend otherwise.

/02 · how the work is shaped

Plan into lanes. Build in parallel. Converge on a verifier that did not write the code.

The build diamond. Agents fan out only where the work is genuinely disjoint, and the human gate sits at the point where a mistake becomes expensive, irreversible or invisible — not at the end by default.

  1. /01

    The harness exists before the code it judges

    Test and contract tasks precede implementation tasks in every phase, because agents cannot verify against a harness that does not exist yet. A team that fans out first and verifies later produces work that must be re-checked by hand — slower than not fanning out at all.

  2. /02

    Worker and checker are never the same agent

    The agent that produced the work may not be the agent that certifies it. Independent checkers, each given a different lens, routinely find real defects that the builder's own green suite had passed. Redundant identical reviewers find less than diverse ones.

  3. /03

    A skipped test is a failure

    Green with a hole in it is worse than red, because nobody investigates green. Our gates grep their own output for skips and exit non-zero. Every artefact we hand over states its coverage boundary and what it does not prove.

  4. /04

    A finding is never waived by the agent that authored the code

    The expensive failure in an operated system is not a missed defect — it is a silently waived one. So the gate sits on the waiver, and a granted waiver is recorded with its justification.

  5. /05

    Each engagement leaves the next one cheaper

    Residue is not a transcript. It is decisions with their reasoning — including the wrong ones, with the correction left visible; evidence that regenerates from a command; skills written down once a method has proved itself; and the project state that is not derivable from the code.

// rule /03, running

Transcript: make adversarial runs six suites. Five pass; one reports SKIP. The gate refuses to close: RED, because a skipped test is a failure and nobody investigates green. The harness is fixed and the gate re-runs; all six suites pass and the gate closes clean with zero skips.

// the only colour on this page marks the slash — and a problem. passing work earns no decoration here, and none in your report either.

A generalist consultancy can copy the org chart. It cannot copy the harness.// why the method is the product

/03 · the exchange

What you receive, and what we need from you.

The artefact test we hold ourselves to: if Pot Labs disappeared, could your own engineers extend this system without guessing? Everything below exists to make the answer yes.

// what you receive

/01

The system

Deployed in your environment, on your infrastructure, under your credentials. Model choice is an engineering decision made per workload, stated in the build, and replaceable.

/02

A constitution

A versioned, ratified list of invariants the system may not violate, authoritative over any instruction given in a session — each article naming its enforcement mechanism in the same table as the rule.

/03

The proving harness

Adversarial suites, a sandbox with simulated counterparties, and CI jobs that fail the build on a violation. Machine checks, not review opinions.

/04

The operating model

The document that tells the next agent — and the next engineer — how work on this system is organised: task classes, fan-out topology, routines, anti-patterns.

/05

Exit criteria as commands

"Done when this suite passes with zero skips and this quickstart step runs clean" — never "done when implemented." A phase closes on its demo, not on its task list.

/06

The uncovered list

What the results do not prove, each limit naming its blocking dependency. A report that cannot fail is a brochure.

// what we need from you

  1. /01

    A named decision-maker who can ratify invariants

    With the authority to say no. This is the hardest dependency and the most common cause of drift. Everything else is logistics.

  2. /02

    The done condition, not the how

    We brief agents at colleague level with a hard exit criterion. We need you to hold the same posture with us.

  3. /03

    The real interface specs — or permission to say we are guessing

    A simulation built on an assumed spec is a guess, and we will label it that way in your own report. Every green is then sandbox-proven, not live-proven, and the report says so.

  4. /04

    The exceptions, not the documented version

    For workflow engagements: one process you already run, with a person who has actually run it. The happy path is easy. The engagement lives or dies on the 3%.

/04 · the boundary

What we do not sell.

Stated as a boundary, because a services page that only says yes is a brochure. Several of these are constraints on us, and we accept them because the alternative is a worse incentive.

  1. /01

    We do not sell headcount

    The commercial unit is an artefact with an exit criterion, not a person on a desk for a quarter. It means an engagement that finishes early finishes, and we invoice less. If you want bodies, we are the wrong firm and will say so on the first call.

  2. /02

    We do not sell a model licence

    We do not own a frontier model and will not pretend a wrapper is one. Anyone selling you their proprietary model when they mean their proprietary prompt is telling you something about the rest of the engagement.

  3. /03

    We do not certify our own work as independent

    Where we built the system, our assurance is internal assurance, and every artefact says so. Where a counterparty needs independence, we will name a third party — including when that costs us the work.

  4. /04

    We do not sell a graph that has not earned its shape

    More agents means more noise and more coordination. Every added node costs context, adds a failure mode, and creates a seam where an assumption can hide. If a graph does not remove fake waiting, separate worker from checker, place the gate where mistakes get expensive, or leave a trail — we run the work inline, charge for less, and tell you why.

  5. /05

    We do not sell green

    A pass report that omits what was not tested is an incomplete report. If that makes our deck look weaker than a competitor's, the competitor's deck is the problem.

  6. /06

    We do not sell a claim we cannot regenerate

    If a number cannot be produced by a command, it does not go in the report — and it does not go on this page. That is why you will find no metrics here, and no client logos: we have not published a figure we could not hand you the command for.

  7. /07

    We do not sell media

    No buying, no placement, no arbitrage on someone else's inventory. Where measurement work touches media, we build the measurement and you keep the buying relationship. A firm that both buys the media and grades it has one client too many.

/05 · start

Bring one process, and the person who actually runs it.

Discovery is fixed in scope and fixed in duration, and it can end the engagement. If the first pass shows the work is interlocked, or that one agent holds it comfortably, we say so and we stop. That is a successful discovery.

// email labs@potstrategy.com
// where Johannesburg, South Africa