TOPIC HUB // 7 ESSAYS

AI Agents

Agents write; humans verify. These essays are about the part nobody demos: what it costs to prove agent output is correct, who owns the orchestration layer, and how to measure autonomy you can actually accept.

-> Proof-Adjusted Autonomy - the metric

OPERATING THESIS

An agent is not autonomous because it finished without interruption. It is autonomous only to the extent that the organization can accept its work with complete, independent and timely proof.

The model is replaceable. The harness, evidence trail, permissions and review capacity usually determine what can run in production.

Proof-Adjusted Autonomy||8 min

Proof-Adjusted Autonomy: The 90% Agent Is a 61.6% Agent

Your agent is 90% autonomous on the demo slide and 61.6% autonomous in the audit. PAA is the metric that explains the difference - and Proof Debt is where it goes.

READ ->
AI Security & Geopolitics||18 min

Token Revocation Is Not an Endpoint

A 200 response from /revoke proves intent, not enforcement. Measure the last path still open, the authority kill graph and the irreversible actions at risk.

READ ->
Future of Work||7 min

Verification Cost Is the New Bottleneck

What's being automated isn't engineering judgment, it's transcription cost. AI collapses creation toward zero while verification cost holds. Engineers move up the stack.

READ ->
Future of Work||7 min

The Unit of Work Is the Agent-Hour

OpenAI's top employees run more than 60 hours of agent work inside a 24-hour day. That isn't overtime. It's a different unit of work: the agent-hour.

READ ->
AI Architecture||6 min

Who Owns Your Harness? The Layer Above the Model

Most companies think they're buying AI. Really they're wiring their whole execution layer around one vendor. That's where lock-in begins.

READ ->
AI Research||7 min

Language World Models: Predict Before You Act

Most agents learn by acting and finding out. Qwen-AgentWorld learns to imagine the world first, then act. The simulator is now a public good.

READ ->
Enterprise AI||7 min

Execution Architecture Beats Model Capability

AI doesn't fail on capability. It fails on the validation structures, decision chains, and error economics companies never built. Execution architecture beats model capability.

READ ->
DIRECT ANSWERS
How should an organization measure AI-agent autonomy?

Measure autonomous completion, complete evidence, independent validation and timely delivery together. Multiplying those gates produces Proof-Adjusted Autonomy.

What usually limits the number of AI agents in production?

Human and independent verification capacity. Agents generate in parallel, while expert review remains scarce and frequently serial.

How do you verify AI agent work?

Require an evidence bundle, then validate it with a mechanism independent from the generator: deterministic tests, replay, a different model family or a human at irreversible boundaries. The agent should never be the only judge of its own output.

What is a good AI agent autonomy metric?

Proof-Adjusted Autonomy is the accepted share of agent work after four gates: autonomous execution, complete evidence, independent validation and on-time delivery. It replaces a demo percentage with a number an operator can estimate from production logs.

OTHER TOPICS

Read these before the feed catches up.

Every essay here is published first at piszczek.pl. Follow along on LinkedIn or Substack.

ALL ESSAYS SUBSTACK