/00 — boot sequence

Hello.

Article

Grok Build CLI Secretly Uploads Your Entire Codebase

July 12, 2026•6 min read
security privacy ai-tools xai grok data-privacy cli

A developer just published a meticulous wire-level teardown of xAI's Grok Build CLI that reveals something alarming: the tool uploads your entire repository to Google Cloud Storage including .env secrets and files you explicitly told it not to read -- and disabling the "Improve the model" toggle does nothing to stop it. The analysis, which went viral on Hacker News with 129 upvotes, is the most detailed independent security audit of a commercial AI coding tool to date.

The findings raise serious questions about data privacy in AI-assisted development tools. As more developers adopt CLI-based coding agents from OpenAI, Anthropic, and xAI, understanding exactly what these tools transmit to their servers is becoming critical.

What Happened

On July 12, 2026, security researcher cereblab published a detailed analysis titled "What xAI's Grok Build CLI Actually Sends to xAI" on GitHub Gist. Using mitmproxy to capture all HTTPS traffic, the researcher documented exactly what the grok CLI binary (version 0.2.93) sends to xAI's servers during normal operation.

The analysis used a throwaway repository with planted "canary" secrets -- unique markers placed in every file so any data appearing in captured traffic could be unambiguously traced back to its source file.

What the Investigation Found

Finding 1: Secrets Transmitted Unredacted

When the Grok Build CLI reads a file, its contents are transmitted to xAI verbatim. The researcher planted a .env file with fake credentials:

API_KEY=CANARY7F3A9-SECRET-should-not-leave DB_PASSWORD=CANARY7F3A9-DBPASS

These secrets appeared in two channels: the live model turn (POST /v1/responses) AND a session_state archive uploaded to POST /v1/storage, which was accepted with HTTP 200. The .env file was transmitted with no redaction whatsoever.

Finding 2: The Whole Repository is Uploaded

This is the most concerning finding. Grok packages and uploads the entire workspace -- every tracked file plus git history -- via POST /v1/storage, completely independent of what the agent reads or is asked to do.

The definitive proof: with the prompt "Reply with exactly: OK. Do not read or open any files.", Grok uploaded the entire repository as a git bundle via POST /v1/storage. Running git clone on the captured bundle recovered src/_probe/never_read_canary.txt -- a file the agent was explicitly told not to open -- with its unique marker verbatim, plus the full git history (4 commits, 47 files).

The finding was replicated on a second, unrelated codebase with the same result.

Finding 3: Multi-Gigabyte Uploads with No Ceiling

The upload mechanism has no practical size limit in the tested range. On a 12 GB repository of never-read random files, Grok uploaded 5.10 GiB across 73 chunks (all HTTP 200) before the capture was stopped mid-stream. Meanwhile, the model-turn channel moved just 192 KB -- a roughly 27,800x ratio proving the upload targets the codebase itself, not what was read.

Repo SizeUpload BehaviorResult
64 MBSingle POST (48 MB)HTTP 200
~600 MBDozens of 7.5 MB chunksAll 200
~3 GBDirect GCS multipart PUT, 50 MB partsAll 200
~12 GB73 chunks at ~75 MB each (5.1 GiB captured)All 200, truncated mid-stream

Finding 4: "Improve the Model" Toggle Does Nothing

The most critical privacy finding: turning off "Improve the model" in settings has no effect on the codebase upload. With the toggle OFF, Grok still uploaded the entire repository as a git bundle, and the server's /v1/settings response still returned:

json

The opt-out governs training, not whether your code is uploaded or stored. The codebase upload to the GCS bucket continues either way.

Where the Data Goes

The storage destination is a Google Cloud Storage bucket named grok-code-session-traces (not AWS S3). This is confirmed by:

  • Binary strings in the grok binary itself
  • A captured metadata.json with per-file destinations like gs://grok-code-session-traces/repo_changes_dedup/v2/...
  • Direct PUT requests to storage.googleapis.com at multi-GB scale

The upload machinery is built into a first-party Rust crate called xai-data-collector, revealed by binary strings such as:

crates/codegen/xai-data-collector/src/gcs.rs crates/codegen/xai-data-collector/src/storage_client.rs crates/xai-grok-shell/src/upload/{gcs,turn,trace,manifest}.rs

Third-party telemetry also fires: POST api.mixpanel.com/track (Mixpanel analytics) and POST grok.com/_data/v1/events.

Why This Matters to Developers

The implications for developers using Grok Build are significant:

  • Secrets exposure: If your project contains .env, secrets.env, or API keys in any file Grok reads, those credentials are transmitted unredacted and persisted to cloud storage
  • Intellectual property risk: Your entire codebase -- including unreleased features, proprietary algorithms, and internal tools -- is uploaded to a third-party cloud bucket
  • Git history exposure: The git bundle includes the full commit history, revealing your project's evolution and potentially exposing previously committed secrets
  • No meaningful opt-out: The documented privacy toggle ("Improve the model") does not affect the upload behavior
  • Disk exhaustion risk: The upload queue can grow to tens of GB on disk during large repo scans, potentially exhausting storage

What Security Teams Should Do

  1. Audit usage: Identify any teams or individuals using the Grok Build CLI in your organization
  2. Review data exposure: Assume your entire repository has been uploaded if the tool has been used
  3. Rotate credentials: Rotate any API keys, database passwords, or secrets that may have been in the workspace
  4. Implement network controls: Block outbound connections to grok-code-session-traces at the network level
  5. Evaluate alternatives: Consider local-only coding agents or tools with verified privacy guarantees
  6. Add to vendor risk assessment: Document this finding in your vendor security reviews

How the Investigation Was Done

The researcher used a straightforward mitmproxy setup:

bash

This technique works because Grok does not certificate-pin against mitmproxy's CA. The analysis was conducted on macOS (Apple Silicon) with grok 0.2.93 in July 2026.

Frequently Asked Questions

Does this mean xAI trains on my code? Not necessarily. Upload and storage do not equal training. That is governed by xAI's policy and your account tier. The investigation measured transmission only, not how xAI uses the data after receipt.

Is this specific to Grok Build, or do other AI coding tools do this too? This analysis covers only the Grok Build CLI. Other tools (GitHub Copilot, Cursor, Claude Code) have different architectures and privacy models. Each should be evaluated independently.

Does the free tier have the same behavior? Yes. The multi-GB upload was demonstrated on a standard consumer account. The git-bundle content upload was confirmed on SuperGrok with "Improve the model" disabled.

Can I stop the uploads? The investigation did not find a setting that disables the upload. Blocking outbound connections to grok-code-session-traces at the network level may prevent the data from leaving your machine.

Does this affect Gitpod or Codespaces users? Any environment where the CLI runs and has outbound internet access is affected. The tool operates the same regardless of the host environment.

Key Takeaways

  • Grok Build CLI uploads your entire repository to Google Cloud Storage, including files you told it not to read
  • .env secrets are transmitted unredacted in both model-turn and storage channels
  • Disabling "Improve the model" has no effect on the upload behavior
  • Multi-gigabyte uploads succeed with no failures and no apparent size limit
  • The mechanism is not documented in the CLI's setup materials
  • The xai-data-collector crate is a first-party Rust component, not a third-party integration

Conclusion

The Grok Build CLI wire-level analysis reveals a significant gap between what developers expect and what actually happens when they use xAI's coding tool. Uploading the entire codebase to cloud storage -- without transparent disclosure and without respecting the privacy toggle -- is a serious transparency failure.

As AI coding tools become essential parts of the developer workflow, this investigation sets a welcome precedent for independent security auditing. Every developer should know what their tools are sending to the cloud.


Sources: Wire-Level Analysis Gist, xAI Privacy Policy, Hacker News Discussion

Automated Transmission

This entry was synthesized and populated dynamically using native API integrations.

Resources & Links