Every answer comes with its sources. Every change comes with a named person's sign-off. Everything lands on a record nobody can quietly edit. scriptorium does real work across your tools, and it can show you proof of all of it.
The AI failures people remember come in three kinds. In each one, something the AI said or did was never checked. scriptorium checks it, in code, every time.
A support bot told customers about a login policy that didn't exist, and they cancelled over it. An airline's chatbot promised a refund the airline didn't offer, and it was held to it.
Receipt: a source. Every claim must cite a record a tool actually fetched. A citation it invented is removed; with nothing to cite, it says so and opens a docs ticket.
An AI agent deleted a production database during an explicit code freeze. No one had signed off on the command.
Receipt: a signature. Every write waits on a card for a listed person, bound by hash to the exact change they saw. The one who asked can't be the one who approves.
The same agent then produced fabricated data and said a rollback was impossible, until it turned out the rollback worked.
Receipt: the record. Every step is written to a hash-chained audit log. Edit one line and verification fails. Nothing it did depends on its own account of it.
One engine and one policy layer run every agent. The Teammate is the general one: it reads widely inside its allow-lists, and each write it proposes waits on a card for a listed approver.
A rule that only lives in a prompt holds as well as the model obeys it. Here each rule is enforced in code, and sabotaged on purpose by a mutation test that must fail when the rule is removed.
A test that still passes with its guardrail deleted proves nothing. So every rule is sabotaged on purpose: pnpm mutate removes it, and its eval must fail. Independent reviews then try to break what's left. Here is some of what they caught before it shipped.
Built in the open: reviewed pull requests, each with a test that fails without it. See them on GitHub →
Files and git are the system of record. No database. An agent is configuration on the same engine: its envelope, its skills and its triggers.
One container and a folder of files; git is the system of record, so there is no database to run. Bring the Anthropic API, Claude on Vertex, or Gemini.
git clone https://github.com/sloweyyy/scriptorium.git && cd scriptorium
cp .env.example .env # a model key, a Slack app, who may approve
docker build -t scriptorium . && docker run --env-file .env -p 8080:8080 scriptorium
pnpm doctor # names anything missing, and how to fix it