All articles Secrets detection

Hardcoded Secrets: Still the Most Common Finding After All These Years

Ariel Ben-David

Hardcoded Secrets Still Common

How Secrets End Up in Code

The paths are consistent across teams, regardless of language or stack. A developer creates a local .env file to test against a staging API. The .gitignore covers the root-level .env but not the one inside a subdirectory added three months later. A quick git add . before a deadline and the credential is in the history. Nobody notices until a rotation audit runs months later.

Other common paths: a developer hardcodes a key to test a specific code path, intends to remove it before committing, commits under time pressure. A configuration template file, config.yaml.example, gets actual credentials filled in during setup, and the filled version ends up committed alongside the template. Test fixtures contain a real API key that was used to capture an integration test response, and the key is embedded in the JSON blob.

Each of these has a different prevention shape. The .env case requires a correct .gitignore and a pre-commit hook. The debugging-artifact case requires regex scanning at commit time. The config template case requires reviewing what files contain credentials before committing. The test fixture case requires scanning test data and test assets, not just application source code. One tool configuration does not cover all four patterns, which is part of why the problem persists despite tooling availability.

What Pre-Commit Hooks Actually Catch

A pre-commit hook running a regex-based scanner catches the obvious patterns: API keys with recognizable format prefixes (AKIA for AWS access keys, sk-live- for Stripe live keys, AIza for Google API keys, ghp_ for GitHub personal access tokens), high-entropy strings in assignment statements, and common JWT formats.

What it misses: a credential split across lines in a config file, a connection string where the password field is not in a recognizable format, a base64-encoded credential, or an internal API key without a known prefix pattern. Entropy-based detection reduces the miss rate on high-entropy strings but introduces false positives on legitimate high-entropy content like encrypted payloads or generated IDs.

A typical gitleaks configuration combining both approaches:

[extend]
useDefault = true

[[rules]]
id = "custom-internal-api-key"
description = "Internal API key pattern"
regex = '''APPKEY[A-Z0-9]{32}'''
tags = ["secret", "internal"]

The useDefault = true pulls in gitleaks' built-in ruleset, which covers around 150 provider patterns. The custom rule extends it for internal key formats. Running this on every push catches commits before they reach the main branch, but only if the CI check is configured to fail the build on a match, not just emit a warning.

Why Git History Creates Persistent Risk

Removing a secret from the current version of a file does not remove it from git history. Running git log -p -- path/to/file shows every version of that file, including commits that contained the credential. The secret remains accessible to anyone with read access to the repository, and if the repository is later made public or cloned by an attacker with temporary access, that history travels with it.

Remediation requires revoking the credential first (always the correct first step, regardless of whether it was used), then optionally rewriting history using git filter-repo or BFG Repo Cleaner. History rewriting is disruptive on active repositories: it changes commit hashes, requiring everyone to re-clone. The disruption scales with how long the secret was in history and how many branches reference that history.

The practical implication is that preventing a secret from reaching the repository costs much less than remediating one after the fact. A commit that a pre-commit hook blocks is trivial to address. A secret that has been in the main branch for three months requires credential rotation, access log audit, and potentially a security incident report depending on your compliance obligations.

The Enforcement Gap

Most teams with a recurring secrets detection problem do not have a detection tooling problem. They have a tool that is installed but not enforced. Pre-commit hooks are opt-in per developer: they exist on one machine, are absent from another, or get removed when a developer reinstalls their environment. The CI check runs but is configured as advisory rather than blocking.

Effective enforcement means a CI check that fails the build on a confirmed secrets match, configured to block merge to the protected branch, with a documented exception process for false positives. The exception process matters: if a false positive blocks a deploy and there is no way to clear it without disabling the check, the check gets disabled. A lightweight exception flow, a review and approval step, not a permanent bypass, maintains enforcement while handling real edge cases.

One pattern that works in practice: run the scanner in two modes. The pre-commit hook is a fast, lower-sensitivity scan that blocks obviously bad patterns immediately. The CI check is a higher-sensitivity scan that runs the full ruleset and blocks merge. The pre-commit hook catches most cases before they reach CI. The CI check is the authoritative gate that cannot be bypassed locally.

The Long Tail: Credentials That Do Not Look Like Credentials

Not all hardcoded secrets have recognizable format prefixes. Database connection strings, SMTP credentials, OAuth client secrets, and internal service passwords are harder to detect with pattern-based rules because the format is not predictable from the key alone.

A database connection string looks like:

DB_URL = "postgresql://appuser:[email protected]:5432/production"

The password Xf9mK2pLnQ has entropy consistent with a credential but may not match a pattern-based rule unless the scanner specifically handles URI credential formats. Entropy scanning catches it; pattern scanning may not. Neither method handles every case, which is why combining pattern matching, entropy analysis, and periodic historical repo scanning produces more reliable coverage than any single approach.

This is not a solvable problem at the tooling level alone. Secrets will be committed as long as developers work under time pressure with credentials available in their local environment. The tooling reduces the rate substantially. The enforcement policy determines whether those reductions compound into an actual program or remain aspirational. The scanner being present is not the same thing as the scanner blocking bad commits.

Tenzai

See findings and fixes together, not just findings.

Free plan available. Connect your first repository in under five minutes.

Start Free Trial

More from the Tenzai blog