Skip to content

Repository files navigation

Code Minification Skill

Slim down source code for AI agents — save tokens during exploration

Quick Start Agent Skills MIT
Stars 中文


An Agent Skill for Claude Code, Codex, opencode, and compatible AI coding agents. It strips comments and formatting noise from source files for read-only exploration while preserving strings, raw strings, template literals, and common regex literals. Zero external dependencies.

✨ Features

  • Zero dependencies — pure Python stdlib, no pip install
  • 13 languages — Python (tokenize), JS/TS, Go, Rust, Java, C/C++, C#, Swift, Ruby, Shell
  • Idempotentminify(minify(x)) == minify(x), safe for tool round-trips
  • Syntax checks when available — Python, Go, JS, Rust, Java, C/C++, Swift, Ruby, and Shell checks are supported when local parsers exist
  • Lexical safety — comment markers inside protected literals are preserved
  • 3-phase eval — built-in evaluate.py: reduction, syntax, LLM comprehension

🚀 Quick Start

# Minify a single file
python3 minify_code.py path/to/file.py

# Keep comments
python3 minify_code.py --keep-comments path/to/file.go

# Pipe from stdin
cat file.ts | python3 minify_code.py --language typescript

# JSON output (for automation)
python3 minify_code.py --json path/to/file.rs

🔧 Installation

Claude Code — clone into your skills directory:

git clone https://github.com/keepkeen/code-minification-skill.git ~/.claude/skills/code-minification

Codex — install by name:

$skill-installer https://github.com/keepkeen/code-minification-skill

Standalone — use it as any Python script:

git clone https://github.com/keepkeen/code-minification-skill.git
alias minify='python3 /path/to/code-minification-skill/minify_code.py'

📊 How It Works

flowchart LR
    A[Source File] --> B[Tokenize / Lexical Scan]
    B --> C[Preserve Protected Spans]
    C --> D[Strip Comments Outside Spans]
    D --> E[Normalize Layout]
    E --> F[Optional Parser Checks]
    F --> G[Minified Output]
    
    style A fill:#e1f5fe
    style G fill:#e8f5e9
    style F fill:#fff3e0
Loading

Python uses stdlib tokenize to preserve indentation semantics.
C-style languages use a single-pass lexical scanner rather than raw regex comment stripping.
Evaluation uses local parsers where available and marks unsupported parser checks as skipped.

🌐 Supported Languages

Extension Language Strategy
.py Python tokenize module — indentation-aware
.js .mjs .cjs JavaScript Lexical comment strip
.ts TypeScript Lexical comment strip
.jsx .tsx React Lexical comment strip
.go Go Lexical comment strip, preserves newlines
.rs Rust Lexical comment strip, raw/nested-comment aware
.java Java Lexical comment strip
.c .h C Lexical comment strip
.cpp .hpp .cc C++ Lexical comment strip, raw string aware
.cs C# Lexical comment strip
.swift Swift Lexical comment strip, multiline string/nested-comment aware
.rb Ruby Lexical hash-comment strip
.sh .bash Shell Collapse blank lines

📈 Evaluation

gantt
    title 3-Phase Evaluation Pipeline
    dateFormat  X
    axisFormat  %s
    
    Phase 1: Reduction Metrics :a1, 0, 1
    Phase 2: Syntax + Idempotency :a2, 1, 1
    Phase 3: LLM Comprehension A/B :a3, 2, 1
Loading
Metric                    Result
──────────────────────────────────────
Average token reduction   ~10–35% typical
Syntax validation         Checked parsers pass; others skip
Idempotency               Expected to pass; verify with evaluate.py
LLM comprehension         Best used for read-only exploration

Run it yourself:

python3 evaluate.py samples/*.py samples/*.go samples/*.js

✅ When to Use

  • Exploring a new codebase — read many files, build mental models faster
  • Large files (>100 lines) — cut token cost by up to 50%
  • Token budget constrained — maximize context window usage
  • Cost-sensitive sessions — fewer tokens = lower API cost

⚠️ When NOT to Use

Scenario Why It Fails
🔴 Compile error debugging Error line numbers mismatch minified output
🔴 Stack trace analysis file:line references become useless
🔴 git diff / code review Diff vs minified view are misaligned
🔴 Unsupported config/DSL files JSON/YAML/TOML/DSLs often need exact formatting or line references

See SKILL.md for the full risk table and anti-patterns.

📁 Project Structure

code-minification/
├── SKILL.md              Skill definition (agent-consumable)
├── minify_code.py        Minifier — pure stdlib, 13 languages
├── evaluate.py           3-phase evaluation pipeline
├── test_minify_code.py   Regression tests for lexical edge cases
├── README.md             This file
├── README.zh-CN.md       Chinese translation
├── LICENSE.txt           MIT license
└── .gitignore

🙏 Credits

Inspired by vix — an AI coding agent with a native Tree-sitter virtual filesystem for code minification. This skill brings the same idea to any agent via a standalone Python tool.

📄 License

MIT — free to use, modify, and distribute.

About

A code minification skill for AI coding agents — reduces token usage while preserving executable semantics across 13 languages.

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages