Is the execution receipt an external contract or an execution byproduct? #5677
ljefford2-cmyk
started this conversation in
General
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
NemoClaw’s recent work around containment, identity, credential isolation, leasing, and auditability is the part of the stack I would most want to build on. One question determines whether that is feasible, and I have not found a clear answer in the docs or release notes.
When a sandbox executes an authorized job, is the resulting audit record / execution receipt specified by what a downstream consumer is entitled to require from it, or by what the execution layer happened to have available to log?
To scope what I mean by “sufficient”: I am thinking about the needs of an independent evaluator rather than the needs of the execution runtime. A completed job may later need to be reviewed by someone outside the executing context. I am not asking whether the runtime itself assesses organizational impact — that would be the execution system reaching into evaluation, which is precisely what should not happen. I am asking whether the record preserves enough that someone else can.
Concretely, the reviewer must be able to reconstruct the authority under which the work was performed, the sequence of actions taken, the procedures and policies applied before execution, the resulting actions and side effects, and enough surrounding context to later assess impacts that exist outside the agent’s immediate scope — compliance exposure, operational consequences, and cross-functional effects the executing system itself may never have observed.
The distinction I care about is between recording a judgment and preserving the material for a judgment. Almost every runtime can tell you what tool was called, what file was read, what an API returned, and whether an action succeeded. The harder question is whether the receipt preserves enough — inputs, authority, actions, side effects — that an external evaluator could later determine whether the action was appropriate, whether it conflicted with another department’s work, or whether it advanced or harmed a larger objective. The runtime is not responsible for knowing those things. It is responsible for not destroying the information that lets someone else reconstruct them.
That is the standard I am using when I ask whether the receipt is a contract or telemetry.
A concrete version of the question: if I am an external auditor holding only the execution receipt for a completed job — with no access to the live sandbox and no execution-layer internals — what am I guaranteed to be able to reconstruct?
For example:
the authority under which the action was taken;
the tools leased and the boundaries they operated within; the inputs that drove each decision point;
the policy decisions applied before execution;
the resulting actions and side effects; enough provenance to replay or independently evaluate the job.
Or is the receipt’s shape effectively “whatever the execution path emitted,” meaning its sufficiency may vary by job, tool, model, or runtime path?
What I am really asking is whether NemoClaw has a stated receipt/audit contract — a stable “this is what the record guarantees to any downstream reader” specification — that independent evaluators can safely build against across releases.
If not, that is also a valid answer. It simply means integrators should treat sandbox output as a raw capture surface and impose their own evidentiary schema at the boundary.
This is not a criticism of the hardening work. The recent secret-boundary and credential-isolation work is exactly the direction I would expect in a serious runtime. I am trying to determine whether the receipt can be treated as a consumer-facing contract for an independent evaluation layer, or whether it should currently be treated as execution-side telemetry whose schema and completeness are still runtime-defined.
All reactions