How we govern
what we build.
Our framework for risk assessment, agent governance, security, privacy, and the commitments we hold ourselves to as we scale. This is not a compliance document. It is an operating manual: the rules we follow when no one is watching, because the receipts will show it either way.
LIVING DOCUMENT · CURRENT PRODUCT STATESCentra is an agent platform built on a simple premise: agents that do real work must produce real proof. Every action is journaled. Every consequential decision waits for a human. Every finished run ends in a signed receipt. These are not features. They are the structural commitments that make Gideon safe to hand work worth doing.
This framework applies the same principles at every stage of the work: authority stays explicit, consequential actions remain gated, and claims stay attached to evidence. Product-specific controls are described only where they have actually been implemented.
What we watch for.
Risk in an agent platform is not a single category. It is a spectrum that ranges from "Gideon spent $5 without asking" to "Gideon was manipulated into exfiltrating data." We identify, assess, and mitigate across this spectrum.
Consequential Actions
Anything that spends money, writes to external systems, or cannot be undone. These are gated behind human approval, always, without exception, regardless of how the request is framed.
Agent Autonomy
Gideon can work independently on non-consequential tasks. As autonomy increases, so do the policy constraints and monitoring around it. We do not grant autonomy we cannot govern.
Data Handling
What Gideon remembers, what it forgets, and who can access it. Memory is scoped per user, sensitive context fades by design, and cross-user leakage is treated as a critical failure.
Misuse and Manipulation
Multi-step instruction injection, context poisoning, and social engineering attempts. We test for these actively and treat every finding as a safeguard improvement signal.
Operational Risk
What happens when a run fails, a tool is unavailable, or a policy is violated. Every failure mode has a defined response: stop, journal, receipt, and notify.
How we assess.
Assessment is not a single checkpoint. It is a multi-layered process that combines internal evaluation, external review, and live monitoring. Each layer catches what the previous one missed.
Pre-deployment evaluation
Before any change reaches operators, it passes through the evaluation areas described in our Safety & Evaluation Report. No safeguard ships untested.
Threat modeling
Every new capability is examined from an adversarial perspective. What could an attacker do with this? What could go wrong? The design is hardened before implementation begins.
Red team testing
Internal and external testers attempt to bypass safeguards, inject instructions, and manipulate Gideon into crossing boundaries. Findings feed directly into the next iteration.
Operator feedback
Operators in the preview can report issues, unexpected behavior, and near-misses through the platform. These are treated as safety signals, not support tickets.
Continuous monitoring
Receipts, walk transcripts, and policy violation logs are monitored for anomalies. Patterns across operators inform safeguard improvements.
How we keep Gideon in bounds.
Governance is not a document. It is a set of enforceable, verifiable mechanisms that constrain what Gideon can do, record what it did, and provide the evidence to prove it.
Usage Policy
A clear, public document that defines what Centra and Gideon can and cannot be used for. It is written to be read by operators, not lawyers. Violations are detectable and enforceable.
Policy Engine
Spend limits, allowed tools, approval paths: defined by each operator, testable before deployment, and enforced by the console. Every violation generates a receipt.
Approval Gate
The structural boundary between autonomous work and human-supervised work. It cannot be disabled, negotiated, or bypassed through conversation.
Audit Trail
Every run, every step, every tool call, every decision is journaled and signed. The trail is tamper-evident and can be verified without an account.
Account Enforcement
Restrictions, suspensions, and bans can be applied in real time when violations are detected. The response is proportional to the severity and the pattern.
How we protect your room.
Every operator gets a private environment. Gideon lives in your room, not a shared space. Security is designed into the architecture, not bolted on after the fact.
Per-user isolation
Each hosted Gideon account has a private agent workspace. Account isolation is an important part of how access is controlled.
Memory scoping
Cortex memory is scoped to the individual. Cross-user retrieval is architecturally prevented, not policy-prevented.
Signed receipts
Every receipt is cryptographically signed. Tampering is detectable. Verification requires no account and no special access.
Two-party controls
Model environments during evaluation are protected by two-party controls with explicit per-user access validation.
Supply chain inspection
Third-party dependencies and integrations are inspected and controlled. Connectors are reviewed before they enter the directory.
Responsible disclosure
We maintain a publicly accessible channel for reporting security vulnerabilities. Reports are reviewed promptly and credited transparently.
What we show, and why.
Transparency is not a section on a website. It is a property of the system. If agents are going to do real work, the people who answer for that work need to see what happened -- not be told it went well.
Methodology documentation
We publish how Centra is built: the team, the pipeline, the role of AI, and the conventions that govern the codebase. Read it at /how-matrix-was-built.
Safety & evaluation report
This hub. What we test, what we found, what we changed. Updated as our practices evolve.
Receipts as public artifacts
Every finished run produces a signed receipt that can be verified by anyone. Proof of work is a file, not a claim.
Walk transcripts
Every run is a readable, ordered, inspectable sequence of steps. A machine journal that cannot be quietly rewritten.
Legal framework
Our terms, privacy policy, AI agent use policy, and governance documents are published and maintained. They are written to be understood, not to obscure.
What we believe, and build accordingly.
We are a small team building a product for a specific kind of user. But the principles behind Centra have implications beyond our product. These are the positions we hold, and the decisions we make because of them.
Operator-first design
Centra is built for people who answer for the work. Density is a feature. Chrome is a bug. We build for the person who has to explain what the agent did, not the person who wants a demo.
Absence as a feature
The most dangerous failures are the ones nobody notices. Centra makes missing steps and unverified claims visible on purpose. This is a design choice with implications beyond our product.
Receipts as accountability
If agents are going to do real work, they need to produce real proof. We are building the infrastructure for agent accountability that we believe should be standard across the industry.
Human-in-the-loop as principle
We do not believe agents should operate without human oversight on consequential work. This is not a limitation we plan to remove. It is a principle we plan to defend.
What happens after it ships.
Safeguards are not static. They evolve as we learn how Gideon is used, where the boundaries are tested, and what the receipts reveal. Post-deployment monitoring is not an afterthought. It is the feedback loop that makes every other layer better.
This is the other half.
The Safety & Evaluation Report covers what we test, what we found, and what we changed. Together, these two documents describe how Centra is built to be safe enough to hand real work.
