How AI Agents Leak Secrets Through Screenshots and Logs

AI agents that take screenshots, read terminal output, or summarise logs can capture API keys, tokens, and customer data. The leak happens when that material leaves the machine: posted to a channel, attached to an issue, sent to a provider, or saved to memory. Prevent it with secret hygiene, output redaction, and limits on where agents can send content.

What agents see

Computer-use agents take screenshots to understand the screen. Coding agents read terminal output, configuration files, environment variables, and logs. Assistants summarise dashboards and error traces. All of these regularly contain sensitive material: API keys printed in a debug log, tokens in an environment dump, customer records in a database console, internal URLs and hostnames, and private messages visible in another window.

Agents capture this material without distinction. A model does not reliably recognise that a string is a secret, and even when it does, it may include it in a summary because it seems relevant.

How the leak happens

Captured material becomes a leak when it crosses a boundary:

  • Posting. An agent with access to chat, issue trackers, or social accounts attaches a screenshot or pastes output to illustrate a problem or report progress.
  • Public artefacts. Screenshots and logs end up in pull requests, public issues, documentation, or demo recordings.
  • Model providers. Content sent as context to a hosted model leaves your environment, subject to the provider's data handling.
  • Persistent memory. Agents that save notes or memories may store secrets in places with weaker protection, to be surfaced later in other contexts.

The common thread is that the agent was given both access to sensitive views and a channel to send content somewhere.

Prevention: keep secrets out of view

The most reliable defence is not letting secrets appear where agents look. Use a secrets manager instead of plain environment files where practical, avoid printing credentials in logs, mask sensitive values in terminals and dashboards, and run agents in environments that hold only the credentials their task requires, scoped narrowly and short-lived. A secret the agent never sees cannot be leaked by it.

Prevention: redact and restrict

  • Redaction. Scan agent outputs, attachments, and memory writes for known secret patterns and sensitive data before they leave the environment, and block or redact matches.
  • Destination limits. Restrict where agents can post or send content. An agent that reads production logs should not also have rights to public channels.
  • Approval for external sharing. Require a human to approve anything an agent posts outside the team, especially images, which are hard to scan reliably.
  • Secret scanning on repositories and trackers, so anything that slips through is caught and rotated quickly.

When a secret leaks

Treat it as an incident. Revoke and rotate the credential immediately rather than only deleting the post, because copies may already exist. Check the credential's usage logs for activity since exposure. Remove the material from the channel, and review why the agent could see the secret and send it where it did. Then fix the access or the destination, not just the instruction.

A checklist for agent deployments

  • List what each agent can see: screens, terminals, logs, files, dashboards, and connected systems.
  • List where each agent can send content: chats, trackers, repositories, social accounts, email, and memory stores.
  • Remove any combination of sensitive visibility and public or external destinations that the task does not need.
  • Add redaction to outputs and memory writes.
  • Require approval for external posts and screenshots.
  • Confirm secret scanning runs on every destination you can control.

Design principle

Give agents the narrowest view and the narrowest output channels their task needs. Retrieval systems can help here: an agent that queries an index of approved, sanitised knowledge, rather than browsing raw systems and screens, sees less sensitive material in the first place.

Frequently asked questions

Can AI coding agents leak API keys?
Yes. Agents read terminal output, configuration, logs, and environment variables, which often contain keys and tokens. If the agent then posts output, attaches screenshots, commits files, or saves memories, those secrets can leave their intended boundary. Keeping secrets out of the agent's view and redacting outputs are the main defences.
What should I do if an agent posted a secret?
Revoke and rotate the credential immediately, since deleting the post does not remove copies. Review the credential's usage logs for unauthorised activity, remove the material from wherever it was posted, and examine why the agent could both see the secret and send it there. Fix that access path.
Are screenshots riskier than text for leaks?
Often, yes. Screenshots capture everything visible, including other windows, notifications, and values the agent was not focused on, and image content is harder to scan automatically than text. Require human approval before agents share screenshots outside the team, and keep sensitive applications off screens agents can capture.
Do AI agents remember secrets they have seen?
Some agents save notes or memories that persist between sessions, and logs of agent sessions may also retain what the agent read. If a secret appeared in context, it can end up in those stores. Scan memory and log writes for secrets, restrict who can read them, and rotate any credential that was captured.
Can redaction tools catch secrets in screenshots?
Partly. Optical character recognition can extract text from images for scanning, but accuracy varies with resolution, fonts, and layout, and context such as customer names or internal details is harder to detect than key patterns. Treat image redaction as a backstop and rely mainly on keeping secrets off screens and requiring approval before sharing.