Claude, ChatGPT, Codex, Cursor, VS Code, Windsurf, Zed, OpenClaw, and other MCP-capable clients enter the same organizational context without being flattened into one UI.
MCP clients · native plugins · web appOrgX
The prompt ends. The company keeps moving.
A wall of busy agent runs, a check that holds, and a receipt you accept. Silent cut.
OrgX lets one person run a fleet of agents from the AI client they already use. The goal, the decisions, the quality bar, and the proof survive every handoff. It lives underneath the clients. It isn't another chat app.
npx @useorgx/wizard@latest setupDetects supported clients, pairs auth, and writes managed MCP configuration.The consequential problem
Agents could do more work than one person could follow.
The models could draft, research, code, and coordinate. The work still scattered across prompts, clients, repos, and whoever happened to remember. One handoff from Claude to Codex could wipe the goal, the decision behind it, and the evidence already gathered.
More autonomy made it worse. Activity went up while the person in charge got less sure about what moved, why, and what needed them.
- Context vanished at client and repository boundaries.
- Output was separated from decisions, constraints, and evidence.
- Human review scaled with activity instead of consequence.
What I saw
The company has to remember why the work exists. The chat can't.
The next agent needs more than a summary. It needs the goal, what done looks like, the decisions and why, the tools, the missing permissions, the budget, and the next useful move. OrgX compiles that into the handoff instead of dumping a transcript.
That turned OrgX from a place you go to orchestrate into infrastructure that meets you where the work already happens.
The decision that changed the system
One work graph across every client. Consequence decides when a person steps in.
The work graph connects goals, initiatives, agents, decisions, artifacts, receipts, cost, and value. The owner's quality bar decides what comes back for review. Publishing, payments, messages, and merges pause at the human boundary. When something fails, the system can retry, narrow the scope, ask a specialist, checkpoint, or stop.
Intent, constraints, decisions, and prior evidence enter together.
Specialists act in the interface best suited to the work.
Consequence determines when human judgment must enter.
The artifact, provenance, and quality outcome become memory.
Independent evidence / reproducible use
Capability is scaling faster than inspectability.
Four systems outside OrgX converge on the same missing layer. The inference is mine; the underlying evidence is not.
OpenAI describes the shift from short chatbot interactions to delegated work that can run for minutes or hours across tools and environments.
OpenAI Economic Research · Jun 2026 ↗02 / ReliabilityCompletion remains probabilistic.METR measures frontier-agent time horizons at both 50% and 80% reliability across more than one hundred software tasks. Capability is rising; a demo is still not a guarantee.
METR Time Horizons · May 2026 ↗03 / EvaluationThe final answer is not enough evidence.Anthropic’s agent-evaluation guidance combines automated evals, production monitoring, transcript review, and human judgment because no single layer catches every failure.
Anthropic Engineering · Jan 2026 ↗04 / PrecedentConsequential automation already travels with attestations.The CNCF-graduated in-toto framework verifies that supply-chain steps happened as planned, under the right authority, without the result being altered in transit.
in-toto · stable specification ↗A work receipt is the narrowest contract that closes this gap.
Taken together, the sources support a bounded claim: when software acts across steps and tools under delegated authority, the result needs inspectable intent, actor, authority, actions, artifacts, evidence, outcome, verification, cost, lineage, and human intervention. They do not prove OrgX is the only answer. Independent emitters, retained users, and paid outcome lift remain open tests.
intentactorauthorityactionsartifactsevidenceoutcomeverificationcostlineageThe claim has a public failure boundary.
On July 27, 2026, the live validator accepted the Codex fixture and rejected the missing-authority fixture with schema.required.
- 30-day uptime
- 99.87%
- Listing score
- 95 / 100
- Evidence date
- 27 Jul 2026
Platform-reported usage proves the server is being called outside the portfolio. It does not yet prove retention, customer outcomes, or revenue. ↗
From install to inspectable handoff in four moves.
- 01Connect the clients already in use.
npx @useorgx/wizard@latest setupThe Wizard detects supported surfaces, pairs auth, writes managed MCP config, and offers companion plugins.
Read the Wizard guide ↗ - 02Make installation prove reachability.
npx @useorgx/wizard@latest doctorA written config is not counted as a working path. Doctor checks the connection and exits non-zero on blocking failures.
Inspect the package ↗ - 03Run a continuity test, not a dashboard tour.
Show me what shipped, who approved it, and the evidence.Ask in one client, then continue in another. The claim: the work graph survives the handoff, not just the transcript.
Open the Agent Amnesia Test ↗ - 04Inspect the returned surfaces.
smithery mcp add useorgx/orgx-mcpUse the independent registry listing, then inspect every MCP App surface in the live widget gallery.
Open the live Widget Gallery ↗
System anatomy / rationale / surfaces
Five mechanisms turn a fleet of agents into accountable company movement.
The difficult part is preserving the cause chain: why the work exists, what the agent knew, which tools it could use, where a person had to decide, what the result proved, and whether the next move is still worth its cost.
Pressure did not decorate the architecture. It determined it.
A new agent receives a task title and reconstructs the company from scattered chats, docs, and repo state.
Compile the goal, definition of done, decisions and why, proof, permissions, confidence, and next action into one context pack.
The next client can continue the work without pretending a transcript is organizational memory.
Streamable HTTP and SSE coexist behind OAuth 2.1, PKCE, and dynamic client registration. Each MCP session is isolated in a Durable Object rather than trusted as a stateless chat request.
Cloudflare Workers · OAuth Provider · Durable Objects · SQLiteA canonical tool grammar moves from bootstrap and search through plan, spawn, decide, write, attach, act, and submit receipt. Zod contracts keep calls structured and interoperable.
Model Context Protocol · Zod · structured resultsInitiatives, workstreams, milestones, tasks, agents, decisions, approvals, artifacts, and outcomes retain the relationships a new session needs in order to continue responsibly.
Next.js · TypeScript · Supabase · React QueryAgent runtimes, queues, sandboxes, and durable workflows execute the work. Trust ladders and consequence-aware gates determine when an operator must enter.
OpenAI Agents SDK · Anthropic Agent SDK · Inngest · E2B · Trigger.devArtifacts return with versions, provenance, evaluation, review state, and a receipt. Proof rooms and embedded widgets make that state legible beyond the dashboard that created it.
MCP Apps · artifact renderers · evaluators · proof roomsThe next agent can resume from decisions, owners, artifacts, approvals, and proof instead of reconstructing a transcript.
Clients keep their native strengths while organizational state remains portable and accountable.
A person only gets pulled in for irreversible or ambiguous calls. Not every agent action.
A reviewer can inspect the artifact, status, provenance, and next action inside the client where the work arrived.
The system becomes tangible through the places people encounter and use it.
Agent desk + chat timeline
Focus, delegation, approvals, tool calls, and outcomes stay attached to the agent's current work.
Live room + processing inspector
Active execution, handoffs, blocked decisions, and run state become legible without pretending raw telemetry is judgment.
Artifact viewer
Code, design, video, data, diffs, marketing work, receipts, and pull requests render on their own terms.
Quality + trust controls
Versioned quality bars remain separate from observed signals; autonomy is bounded by consequence.
Proof rooms
Selected outcomes become durable, shareable capsules instead of screenshots without provenance.
MCP widget system
Initiative pulse, morning brief, decisions, search, status, and task surfaces bring the work graph into AI clients.
Client + host plugins
OpenClaw and other client bridges inherit host strengths while adding shared organizational memory.
Benchmark + evaluation
Judged criteria, receipts, and publication artifacts turn quality claims into a repeatable evidence system.
Next.js App Router · React · TypeScript · React Query · Xyflow
Supabase · PostgreSQL · Clerk · OAuth 2.1 · PKCE
OpenAI Agents SDK · Anthropic Agent SDK · Inngest · E2B · Trigger.dev
Model Context Protocol · MCP Apps · Cloudflare Workers · Durable Objects · Zod
Sentry · PostHog · OpenTelemetry · Upstash · Stripe
The tools behind the decisions.
Select a tool to see the role it plays in this system.
- Protocol / Cross-client continuity
MCP
One portable contract exposes organizational memory and governed actions inside Claude, ChatGPT, Cursor, Codex, OpenCode, OpenClaw, and other MCP hosts.
These are working MCP and operator surfaces, not concept renders.

Organizational health, the blocking boundary, workstream progress, and a live continuation action in one embedded surface.

The widget distinguishes current focus, blocked work, review, artifacts, and progress instead of compressing everything into online or offline.

The next session starts from the company's actual state and the calls waiting on you. Not a blank prompt.

The information needed to decide, the consequence, and the approval action stay together inside the client.

A new session can recover the relevant decision and artifact without reading the entire organizational transcript.

A plan with owners, boundaries, and a next action. Not a paragraph called a plan.
Authentic proof

Output, provenance, quality, and review state in one place. Nothing reconstructed after the fact.

The interface privileges current focus, the next consequential boundary, and grounded history over raw activity.

Decision, trust, and work events remain attributable in the same timeline instead of dissolving into logs.

The bar is explicit, versioned, and task-specific; observed signals stay visible without masquerading as the standard.
What changed in my operating model
The product is the quality of the judgment it makes possible.
OrgX moved me from “automate the workflow” to “design the cause chain.” Every action that matters should keep its context, show its boundary, and come back with proof the next decision can use.
Autonomy is useful. Continuity is what makes it compound.
- Escalate consequence, not mere activity.
- Keep the score-bearing quality bar separate from runtime signals.
- Treat the receipt as a first-class product surface.