Cost per verified outcome
I count inference, tool calls, retries, review time and failed work. Token price alone does not tell me what the business paid.
I am CTO at Archdesk. I work on AI systems used inside enterprise workflows, where a wrong answer can delay a project, trigger rework or create liability. My focus is the layer around the model: cost, data, permissions, verification and recovery.
A model can generate an answer. The company still needs to know what data it used, what it changed, who can approve the result and how to recover when it is wrong. I treat that complete path as the system.
Generation is easy to demonstrate. The harder job is deciding which machine-produced work the organization can accept without repeating the work by hand. NULLIUS IN VERBA // TAKE NOBODY'S WORD FOR IT
I use four ledgers when reviewing an AI deployment. Each one needs a named owner and an answer that can survive an incident review.
I count inference, tool calls, retries, review time and failed work. Token price alone does not tell me what the business paid.
I keep business context, permissions, evaluations, observability and recovery outside the model API, so the model can be replaced.
The model cannot be the only judge of its own output. I use tests, replay, a different model family or a human at the irreversible boundary.
Before production, every consequential action needs a named owner, an evidence record and a recovery path. If nobody owns the failure, the feature is not ready.
I named these measures because model benchmarks did not answer the questions that kept appearing in architecture and budget reviews.
The share of completed work that is autonomous, independently checked and delivered before the decision deadline.
A way to compare AI systems by useful work per unit of energy, instead of treating raw capability as the only score.
The control plane around the model: context, tools, permissions, memory, evidence, evaluation and recovery.
The irreversible work an AI system can still accept while a revocation decision propagates through gateways, workers and regions.
Energy per accepted result, including failed attempts and retries - the production denominator behind the Joule Wars.
Definitions, formulas, DOI records and attribution guidance live in the canonical concept registry.
I did not arrive at this view from model benchmarks. It came from building systems that move money, monitor structures, interpret regulations and run enterprise workflows.
I lead technology at Archdesk, where software sits inside construction, manufacturing and engineering workflows. Access control, auditability and recovery are product requirements.
I founded Lextron.ai, Inclify, Rejsomat.pl and Robotero. The products covered regulatory research, structural monitoring, travel marketplaces and algorithmic trading.
Security research from 2004-2012, including responsible disclosure of critical flaws in Sun Microsystems' sun.com. Read the record.
Experiments, operating models and arguments that I am willing to publish with the numbers attached.
A draft can be regenerated. A payment, production deploy or message to a customer needs a different release gate. I design backward from that boundary.
If the same model writes and grades the answer, both errors are correlated. The check needs another mechanism.
I include retries, review time and human exceptions. A cheap token can still produce expensive work.
Models rotate. Context, permissions, evaluations and recovery logic should remain under the operator's control.
If a human has to reconstruct the work before accepting it, the work was not autonomous in any useful operating sense.
Piszczek, M. AI CTO - Enterprise AI Systems, Proof and Economics. piszczek.pl. https://piszczek.pl/ai-cto
AI deployment, systems architecture, C-level advisory, podcasts and keynotes.