An internal TELOS run, slowed open

Watch one check happen.

This is an internal TELOS run: our own framework, watching our own agent re-run nine of our published safety benchmark results, every step observed. Below is one of those steps: the real sub-second moment at the before-tool-call hook, stretched open so you can watch each part of the check land. Every number on screen comes from the record.

The before-tool-call hook · internal TELOS run · cut: the clean check

The action

incoming

The checks
purpose
0.000
scope
0.000
tool
0.000

Instrument reads, not production-calibrated grades. In passive mode they inform human review, never control execution.

  • metered external APIs
  • push, deploy, publish
  • engine, scoring, keys
  • writes outside run dirs
  • samples as full results
  • raw harmful content

The record
0:00 / 0:42

Observation only: TELOS watches and records at this hook. It does not block, and it cannot.

TWO WAYS TO READ THIS

Get the result in two minutes, or inspect the whole record step by step.

01 · The assignment

First, the job was written down.

Before the agent ran a single step, a person wrote its purpose down and locked it for the run. It holds still while the conversation moves. That fixed point is what every step of the work gets measured against.

THE DECLARED PURPOSE, VERBATIM FROM THE RECORD

frozen manifest

"signed" in this quote means the record's own demonstration key, not a production signature. Production signing is not claimed.

The job came with hard lines.

Six of them, declared up front, quoted here as written:

    TELOS ran in passive observation only. It could read every step. It could not touch one. That is the design, not a limitation: the agent does the work, TELOS shows the work.

    Before real work started, the run tested its own watcher. Step 1 exercised the purpose check against one allowed action and a set of forbidden test fixtures, a fake push and a fake production key among them, and confirmed the reads came back. Then the work began.

    02 · The run, observed

    Then the agent worked. TELOS watched.

    Thirty steps: loading the frozen job, fetching public datasets, hashing them, running nine benchmarks, reading the results, assembling the record. Every step got the same treatment you watched above. Step through the real record here.

    Click a link in the chain, or use the arrow keys.

    How one check works.

    Each check is sub-second, and simple to say. The declared purpose sits at a fixed point in a space of meaning. Every action lands somewhere in that same space. How close it lands to the purpose is the read. Slow that down and it looks like this:

    A fixed purpose anchor with three actions at increasing distance from it the declared purpose lands close: within purpose further out: question raised far from purpose: surfaced for review

    The anchor never moves. The actions land where they land. The distance is the read, and every read is written into the record.

    For the reader who wants the math

    The purpose is embedded once, at declaration, into a fixed vector. Each action is embedded into the same space as it happens. The read is the cosine between the two: the dot product of the vectors over the product of their lengths. A fixed anchor, a moving action, one angle.

    fidelity(action) = cos θ = ( E(action) · E(purpose) ) / ( ‖E(action)‖ ‖E(purpose)‖ )

    ComingThe runnable version of this page's whole check, as a repository you can clone, publishes with the capsule.

    03 · The shown work

    Seven reproduced. One drift. One boundary held.

    Nine benchmark rows, the published numbers next to what this run measured. Seven matched. One drifted under a different scorer. One hit a declared hard line and could not run. The record keeps all three, differences and all.

    A record that can admit a drift and a boundary is the point.

    Published benchmark results compared with this run's results
    BenchmarkItemsPublishedThis runDisposition

    The drift and the boundary are in the table above, not footnoted away. Each row traces back to its raw record in the run explorer.

    04 · Check it yourself

    Check the record. Right here.

    Your browser recomputes the fingerprint of all thirty steps and follows every link down to the pinned head. The math runs on your machine, with no server behind it. Then break the record on purpose and watch it tell on itself.

    The record is loaded. Thirty steps, untouched.

    Purpose In. Proof Out.

    A record keeper for your agentic work.

    That was a real run. Want this for your agents? Talk to the team that built it.

    Email the TELOS team about one workflow ComingSDK coming, not yet available

    The SDK publishes with the capsule. Until then, the mailbox is real and the record above is checkable.

    Purpose In. Proof Out.

    TELOS makes AI-agent work visible and human-accountable.

    info@telos-labs.ai
    TELOS AI Labs Inc. · 2026