A cloud-security credential scanner that goes beyond simple regex matching. CredScan layers pattern matching, Shannon entropy analysis, and context-aware scoring to detect hardcoded secrets across source code, Infrastructure as Code, CI/CD pipelines, Dockerfiles, git commit history, and web endpoints. It can verify which keys are still live and correlate passwords against known breaches.
Try it live: credscan.tolubanji.com. Paste a config or upload files in the browser, nothing to install. The hosted demo runs an upload-only sandbox; no path scanning, no data kept.
| Surface | Command | What you get |
|---|---|---|
| Online | credscan.tolubanji.com | Upload or paste in the browser; sandboxed, masked, nothing stored |
| CLI | pip install -e . then credscan -p . |
Full scanner: paths, git history, web, verification, every report format |
| Local GUI | pip install -e ".[gui,aws]" then credscan-gui |
The full tool in a browser on your machine (path scan + history + validation) |
| Docker (public) | docker run -p 8000:8000 credscan-gui |
The hardened upload-only image, to self-host |
| Docker (local) | docker run -p 127.0.0.1:8000:8000 -v "$PWD:/scan:ro" credscan-gui-local |
Full power in a container, loopback only |
For a walkthrough of every surface and feature in the same terminal style, see the in-app guide at credscan.tolubanji.com/guide. For hosting your own instance, see docs/DEPLOY.md.
git clone https://github.com/ToluGIT/credscan.git
cd credscan
pip install -e .Requires Python 3.9+. Optional extras: pip install -e ".[aws]" for live AWS
key validation, ".[reports]" for Excel/PDF output.
Most credential scanners apply a regex pattern and report a match. CredScan runs every finding through four layers before reporting it:
- Pattern matching: 15+ categories of regex rules covering cloud providers, payment processors, messaging services, database URIs, cryptographic material, and generic secret patterns
- Shannon entropy analysis: high-randomness strings are flagged even without a keyword match; thresholds are tuned per token type (base64 keys have different entropy profiles than hex secrets or JWTs)
- Context analysis: the surrounding lines are examined to assess whether the credential is in production config, test code, documentation, or an example file; confidence is adjusted accordingly
- Confidence scoring: a weighted score combining pattern strength, entropy, context, and technology signals produces a final confidence value; low-confidence findings are filtered before output
The aim is high signal with minimal noise, so findings reported at default settings are worth investigating. Two optional, read-only passes go further: live verification confirms which keys are still active (AWS, GitHub, GCP, Slack, Stripe, OpenAI, Anthropic, npm), and breach correlation checks passwords against the HaveIBeenPwned corpus using k-anonymity, so the value never leaves the machine.
Accuracy claims are backed by a reproducible benchmark, not asserted. Against the bundled labeled corpus (17 planted secrets across 6 files, plus 4 clean files of decoys: env references, placeholders, hashes, UUIDs, base64 config):
| Metric | Score |
|---|---|
| Precision | 1.00 |
| Recall | 1.00 |
| F1 | 1.00 |
PYTHONPATH=src python benchmarks/run.pyThis is a regression suite, not an independent benchmark. The corpus and
several detector fixes were authored together, so the perfect score records that
those fixes behave as intended on representative inputs; it is not evidence of
generalization, and it is not a production precision figure. Its purpose is to
fail CI (--fail-under-f1 0.90) if a change degrades detection. A neutral
cross-tool comparison on a larger labeled dataset is future work. See
benchmarks/README.md for the full design notes.
Test coverage is currently 27% (pytest --cov); the detection pipeline,
parsers, and SARIF output are well covered, while the git-history, web, and
binary subsystems are not yet unit-tested. See SECURITY.md for
the tool's own threat model.
Source types
| Source | How |
|---|---|
| Files and directories | Recursive scan; parsers selected per file type |
| Git commit history | Every commit diff is scanned; secrets deleted from code still exist in history |
| Git staged files | Pre-commit hook mode blocks or warns before a commit lands |
| Web endpoints | HTTP fetch with optional crawling |
File types with dedicated parsers
| Parser | Handles |
|---|---|
| IaCParser | Terraform .tf/.tfvars, CloudFormation YAML/JSON |
| CICDParser | GitHub Actions workflows, GitLab CI, CircleCI, Jenkinsfiles |
| DockerParser | Dockerfiles (ENV/ARG instructions), Docker image tarballs |
| CodeParser | Python, JavaScript/TypeScript, Java, Go, Ruby, C/C++, C#, PHP, Kotlin, Swift |
| JSONParser / YAMLParser | Config files, API specs, package files |
| BinaryParser | ZIP, TAR, JAR, WAR, APK, IPA archives |
| WebScanner | HTML pages, JavaScript bundles, API responses |
Credentials and tokens
- AWS Access Key IDs (
AKIA...) and Secret Access Keys - GCP service account keys and API keys (
AIzaSy...) - Azure connection strings and access keys
- Stripe secret/publishable keys, PayPal Braintree, Square
- GitHub, GitLab, and Bitbucket personal access tokens
- Slack API tokens and webhook URLs
- Twilio, SendGrid, Mailgun, Postmark API keys
- OpenAI, Anthropic, and Hugging Face API keys
- Database connection strings: PostgreSQL, MySQL, MongoDB, Redis
- Generic password assignments, JWT signing secrets, OAuth client secrets
Cryptographic material
CredScan detects private keys and certificates committed directly to repositories. All are flagged at critical severity:
- RSA, DSA, EC, and OPENSSH private keys (
-----BEGIN RSA PRIVATE KEY-----) - PGP private key blocks (
-----BEGIN PGP PRIVATE KEY BLOCK-----) - PKCS#12 and PFX certificate bundles (references to
.p12,.pfx,.pem,.keyfiles) - X.509 certificates (
-----BEGIN CERTIFICATE-----)
Infrastructure as Code
- Hardcoded provider credentials in Terraform (
access_key,secret_keyinprovider "aws"blocks) - Passwords in variable defaults and resource properties
- CloudFormation parameters without
NoEcho: true - Secrets hardcoded in CI/CD
env:blocks instead of referenced from a secrets store
A credential that was committed and later deleted still exists in git history. Every clone of the repository has it. CredScan walks every commit diff in the specified range, applying the full detection pipeline to each change set.
# Scan the entire history of the current branch
credscan --scan-history
# Scan the last 200 commits on main
credscan --scan-history --branch main --max-commits 200
# Scan a specific time window
credscan --scan-history --since "6 months ago" --until "1 week ago"
# Scan history and export findings as SARIF
credscan --scan-history -o sarif -d ./reportsEach finding includes the commit hash, author, timestamp, and the file and line where the secret appeared, giving you exactly what you need to assess exposure and determine when rotation is required.
# Scan the current directory
credscan
# Scan a specific path, group output by severity
credscan -p ./src --group-by-severity
# Scan Terraform and CloudFormation files
credscan -p ./infra -o json,sarif -d ./reports
# Scan a web endpoint
credscan --url https://example.com/static/app.js
# Validate any AWS keys found (calls sts:GetCallerIdentity, read-only)
credscan -p . --validate-aws
# Verify tokens against provider endpoints (GitHub, Slack, OpenAI, and more)
credscan -p . --verify
# Correlate passwords against known breaches (HIBP, k-anonymity)
credscan -p . --check-breaches
# Tune confidence threshold to reduce false positives
credscan -p . --min-confidence 0.6 --entropy-threshold 4.5Exit codes: 0 = clean · 1 = credentials found · 2 = argument error
A terminal-styled web interface drives scans and explores findings for anyone
who would rather not use the CLI. It runs the same engine; the API masks every
value, so no raw secret leaves the server. A hosted upload-only instance is at
credscan.tolubanji.com, with an in-app
guide at /guide.
pip install -e ".[gui,aws]" # aws extra adds boto3 for live key validation
credscan-gui # local mode, http://127.0.0.1:8000Launch a scan, watch detector output stream in live, then filter the findings report by severity, expand a finding for its context and remediation, and suppress false positives into a baseline.
The backend is a thin FastAPI layer (credscan/gui/server.py) wrapping the
existing engine; the frontend is a single static page built to the terminal
design system. Findings are returned masked (AKIA...MPLE), never raw.
- Local (default): the full tool in the browser. Scans server-local paths,
walks git history, and runs live AWS key validation (with the
awsextra). For running on your own machine. - Public (
credscan-gui --public, orCREDSCAN_PUBLIC=1): a hardened mode for hosting on the open internet. Filesystem path scanning, git-history, and live validation are disabled; the only input is uploaded files or pasted text, scanned in a per-request sandbox directory and deleted immediately. Hard limits bound every request (2 MB, 200 files, 30 s, rate-limited). This exists because a publicly reachable path scanner would let any visitor read the host's own filesystem, and live validation would make the server a credential-checking oracle.
The two modes ship as two separate images on purpose, so the public one is safe
by construction rather than by a flag: the public image has no path access and
does not even install boto3, so it cannot scan the host or validate keys no
matter how it is configured.
Dockerfile.gui runs the hardened public mode as a non-root user, upload-only:
docker build -f Dockerfile.gui -t credscan-gui .
docker run -p 8000:8000 credscan-gui # public mode, upload-onlyA fly.toml is included for a one-command deploy to Fly.io (fly deploy); any
container host (Render, Railway) works the same way.
To get the complete GUI (path scan + git-history + AWS validation) in a
container on your own machine, use Dockerfile.gui.local. It runs local mode
and includes boto3. Because local mode scans real paths and can validate
keys, publish the port to loopback only so it is unreachable from the network:
docker build -f Dockerfile.gui.local -t credscan-gui-local .
docker run --rm -p 127.0.0.1:8000:8000 \
-v "$PWD:/scan:ro" credscan-gui-local # scan /scan in the GUIMount the directory to scan at /scan. For live AWS validation, also mount your
credentials read-only (-v "$HOME/.aws:/home/scanner/.aws:ro") and toggle
validate-live in the GUI options. Do not expose this image on the open
internet; that is what the public image above is for.
Reports are generated with --output and saved to --output-dir:
| Format | Use case |
|---|---|
console |
Default; colored, with confidence scores and context |
json |
Audit log; full finding detail plus remediation guidance |
sarif |
GitHub Code Scanning, VS Code, and other SARIF-compatible tools |
html |
Shareable report; values are masked and HTML-escaped |
excel / csv |
Spreadsheet-based triage |
pdf |
Printable report |
compliance |
CSV pivoted by control framework (CWE, NIST 800-53, PCI-DSS v4.0, OWASP ASVS, SOC 2, ISO 27001), with a finding ID, verification status, and remediation |
Secret values are masked in all human-readable output (AKIA...MPLE) and full values are only present in the JSON audit log. The HTML report is generated with proper escaping so content from scanned files cannot execute as code in the browser. The SARIF output validates against the official SARIF 2.1.0 schema, carries CWE tags, and uses stable partialFingerprints for dedup across runs.
For day-to-day use, a few common workflows:
credscan -p . --create-baseline .credscan-baseline.json # save current findings as the baseline
credscan -p . --baseline-file .credscan-baseline.json # suppress everything in it on future scans
credscan --staged # scan only git-staged files (pre-commit)
credscan --diff origin/main # scan only what changed vs a base branch (CI)
credscan --install-hook # install the git pre-commit hookDiff mode reads the changed-file list from git, so a typical commit scans in well
under a second regardless of repository size. Run credscan --help for the full
flag reference, or see the in-app guide.
CredScan ships an official GitHub Action. It produces SARIF that uploads to the repository's Security tab, and fails the job when credentials are found:
# .github/workflows/secrets.yml
name: Secret Scan
on: [push, pull_request]
jobs:
credscan:
runs-on: ubuntu-latest
permissions:
security-events: write # required to upload SARIF
steps:
- uses: actions/checkout@v4
- name: Run CredScan
id: scan
uses: ToluGIT/credscan@v1
with:
path: .
min-confidence: "0.5"
- name: Upload SARIF
if: always()
uses: github/codeql-action/upload-sarif@v3
with:
sarif_file: ${{ steps.scan.outputs.sarif-file }}Or run the CLI directly in any pipeline (exit code 1 means findings):
pip install credscan
credscan -p . --no-color -o sarif -d ./reports --min-confidence 0.5A pre-built Docker image is also available:
docker run --rm -v "$PWD:/scan" ghcr.io/tolugit/credscan -p /scanFor repeatable scans, settings (scan path, exclude patterns, thresholds, output
formats, baseline) can live in a config.yaml loaded with credscan --config config.yaml.
CredScan is a detection aid. It will not catch every possible credential exposure and is not a substitute for proper secret management (AWS Secrets Manager, HashiCorp Vault, GCP Secret Manager, etc.). Any credential it finds should be rotated immediately. Detection confirms exposure, not just risk.


