Urgent.News

What's breaking now, across thousands of outlets.

AI

Audit Logs for AI Coding Agents: What to Record, Where to Collect It, and What It Proves

When a coding agent does something surprising — deletes a directory, pushes to the wrong branch, reads a file it shouldn't have — the first question is always the same: what exactly happened, and what allowed it? Shell history and chat transcripts answer that badly. An audit log for AI coding agents answers it well, if you decide up front what to record. This post covers the fields worth…

When a coding agent behaves unexpectedly, the first question is always the same: what happened exactly, and what allowed it to do so? Shell history and chat transcripts don't provide a clear answer. An audit log for AI coding agents does, if you decide in advance what to record. This post explains which fields to capture, the difference between logging decisions and logging outcomes, where to collect events in common agents, and what a tamper-evident log does and doesn't prove.

An agent audit log should answer two main questions during an incident or review: who acted, and what did it try to do. The log should capture details such as the agent, session, person, command, tool, target, decision, reason, outcome, surrounding calls, and any additional context. If the log can't answer these questions without re-reading a configuration file that may have changed, it's not an audit log but a transcript.

A practical record per tool call should include fields like decision_id, run_id, agent, human principal, tool and normalized action, canonical resource, verdict and rule, explanation (reason), risk level, approval ID and approver, which rules were evaluated and which matched, and timestamps with policy versions. Sensitive information, like raw secrets and full credentials, should be left out and replaced with a hash of arguments.

Two common design mistakes to avoid are conflating authorization decisions with execution outcomes and not recording enough details. Using hooks around tool use, like PreToolUse and PostExecution, can help record both the authorization decision and the outcome. Normalizing events into one schema makes it easier for auditors to analyze the data.

To ensure tamper-evidence, use hash chaining, where each record includes the hash of the previous one. This helps detect in-place edits. However, it doesn't prevent deletion or rewriting of the entire history unless you compare against a trusted external checkpoint, like a separate log store or a ticket. Always verify the record count and check for missing logs.

Regularly ship logs off the machine and apply the same retention and access controls as you use for CI logs. Redact sensitive data at the collector if arguments might contain secrets. Here's an example of what a governed decision looks like in Cirvix AgentControl:

```json

{

"decision_id": "dec_mv3wca7f2",

"agent": "pr-triage",

"tool": "read_file",

"action": "fs.read",

"resource": "/tmp/auditdemo/.env.production",

"verdict": "deny",

"rule": "deny-dotenv-read",

"reason": "Reading .env files is denied outside an approved secrets flow.",

"risk": "critical",

"risk_signals": ["credential-access", "read-only-tool"],

"considered": [

{

"rule": "deny-dotenv-read",

"effect": "forbid",

"matched": true

}

]

}

```

You can verify the log using the cirvix tool:

```

cirvix logs --last 5

cirvix why dec_mv3wca7f2

cirvix audit verify --file .cirvix/audit.jsonl --json

```

By including the head hash and record count somewhere the agent can't write, you create an external checkpoint to ensure the integrity of the audit log.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Building a Full RAG Pipeline in Spring Boot — From Ingest to Grounded Answer

Previous articles covered the individual pieces of RAG. You now know what embeddings are, how to chunk documents, and how Pinecone stores and searches vectors.

  • Ingest flow processes raw text into chunks, vectors, and stores them in Pinecone
  • Query flow retrieves top 5 most similar chunks for each user message
  • IngestService manages Ingest flow, deleting existing chunks and adding new ones to vector store

More from Sunday 11 October →