Intelligent Context Compression for the AI Agent Era
Single binary · Zero dependencies · Up to 70% token savings · MCP Native · Docker <10MB
- 💸 The Problem
- 🚀 Quick Start
- 🔥 Killer Features
- 📖 CLI Reference
- 📚 Go Library API
- 🌐 HTTP Proxy Guide
- 🔌 Integrations
- 🚢 Deployment
- 🧠 How It Works
- 📦 Content Types
- 📊 Real-World Performance
- 🔧 Development
- 🤝 Contributing
Every token you send to an LLM costs money. Agent workflows amplify this — tool outputs, logs, RAG snippets, search results, and conversation history pile up fast. A single agent run can easily burn 50,000+ tokens in context alone.
Headroom Go compresses everything your agent reads before it hits the LLM — slashing token costs by up to 70% while preserving semantic accuracy. It's a production-grade Go port of headroom, purpose-built for the AI agent era.
The math is simple: If you spend $1,000/month on LLM API calls, Headroom Go can save you $700/month.
# One-liner (Linux / macOS)
curl -sSL https://raw.githubusercontent.com/superops-team/headroom-go/main/install.sh | bash
# Go install
go install github.com/superops-team/headroom-go/cmd/headroom@latest
# Docker
docker pull ghcr.io/superops-team/headroom-go:latest
# Homebrew (macOS)
brew tap superops-team/headroom && brew install headroom# Pipe anything through it
cat huge_log.txt | headroom compress --stats
# → original_tokens=12500 compressed_tokens=3750 savings_pct=70.0%
# Aggressive JSON compression
echo '{"items":[1,2,3,4,5],"debug":null}' | headroom compress --aggressiveness 0.8
# Transparent OpenAI proxy
headroom proxy --port 8080import headroom "github.com/superops-team/headroom-go"
messages := []headroom.Message{{Role: "user", Content: "ERROR failed\nINFO retry\nINFO retry\n"}}
result, _ := headroom.Compress(messages, headroom.DefaultOptions())
fmt.Printf("Saved %.0f%% tokens\n", result.Savings*100)headroom mcp serve提供 4 个 MCP 工具:headroom_compress、headroom_retrieve、headroom_stats、headroom_read。Claude Code 配置:
{ "mcpServers": { "headroom": { "command": "headroom", "args": ["mcp", "serve"] } } }headroom wrap claude # 打印配置指令
headroom wrap claude --apply # 自动修改 ~/.claude/settings.json
headroom wrap codex --apply # 自动修改 ~/.codex/config.yaml
headroom wrap generic # 通用配置指南启动本地代理 + 自动配置 IDE,Ctrl+C 恢复原配置。
headroom proxy --port 8787POST /v1/chat/completions— 压缩后转发GET /healthz— 健康检查- ✅ SSE 流式响应(
stream:true) - 自动转发
Authorization/X-Request-ID - 50MB 请求体限制,60s 超时,优雅关闭
压缩后可检索原始内容:
store := headroom.NewCCR(headroom.CCRConfig{TTL: 24 * time.Hour})
id := store.Store("original", "compressed", headroom.KindText)
original, _ := store.Retrieve(id)CacheAligner 为输出添加稳定前缀,提升 LLM 提供方 KV Cache 命中率。
Pipeline 模式下保护 <thinking>、<tool_call> 等 XML 标签不被压缩破坏。
registry := headroom.NewCompressorRegistry()
registry.Register(headroom.NewCompressorFunc(headroom.KindText,
func(content string, opts headroom.Options) (string, error) {
return strings.ReplaceAll(content, "secret", "[redacted]"), nil
},
))headroom <command> [flags]| Command | Description |
|---|---|
compress |
压缩 stdin 或文件 |
proxy |
启动 HTTP 代理(OpenAI 兼容,支持流式) |
mcp serve |
启动 MCP Server(stdio 模式,4 工具) |
wrap <agent> |
启动代理 + 配置 IDE(claude/codex/copilot/generic) |
version |
打印版本 |
| Flag | Type | Default | Description |
|---|---|---|---|
--aggressiveness |
float | 0.5 |
压缩强度 0.0-1.0 |
--no-reversible |
bool | false |
关闭可逆压缩 |
--no-align |
bool | false |
关闭前缀对齐 |
--tokenizer-backend |
string | "" |
fallback / tiktoken / huggingface |
--token-budget |
int | 0 |
Pipeline 目标 token 数 |
--enable-pipeline |
bool | false |
启用 Pipeline 模式 |
--query |
string | "" |
diff/search 评分查询词 |
--input |
string | "" |
输入文件(默认 stdin) |
--output |
string | "" |
输出文件(默认 stdout) |
--stats |
bool | false |
打印 token 统计 |
| Flag | Type | Default | Description |
|---|---|---|---|
--port |
int | 8787 |
监听端口 |
--upstream |
string | https://api.openai.com/v1 |
上游 API URL |
--aggressiveness |
float | 0.5 |
压缩强度 |
--no-reversible |
bool | false |
关闭可逆压缩 |
--enable-pipeline |
bool | false |
Pipeline 模式 |
--token-budget |
int | 0 |
目标 token 数 |
环境变量:HEADROOM_API_KEY — 当请求无 Authorization 头时的上游 Bearer token。
headroom wrap <claude|codex|copilot|generic> [--apply] [--port=18787]headroom mcp serve # 启动 MCP ServerModule: github.com/superops-team/headroom-go
func DefaultOptions() Options
func Compress(messages []Message, opts Options) (*Result, error)
func CompressString(content string, opts Options) (string, error)
func NewCompressionEngine(opts Options) (*CompressionEngine, []Warning)
func NewDefaultPipeline() *Pipelinetype Message struct { Role, Content string; Name string }
type Options struct {
Aggressiveness float64 // 0.0-1.0
Reversible bool // 可逆压缩
AlignPrefix bool // KV Cache 对齐
TokenLimit int // 跳过阈值
TokenizerConfig TokenizerConfig
TokenBudget int // Pipeline 目标
Query string // 相关性评分
EnablePipeline bool // Pipeline 模式
Observer Observer // 步骤回调
}
type Result struct {
Messages []Message
CompressedTokens, OriginalTokens int
Savings float64
Warnings []Warning
Steps []CompressionStep
}POST /v1/chat/completions— 压缩后转发(支持stream:trueSSE 流式)GET /healthz—{"status":"ok","version":"v0.8.0","uptime":"..."}
HEADROOM_API_KEY="$OPENAI_API_KEY" headroom proxy --port 8787 &
curl http://localhost:8787/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"hello"}]}'curl http://localhost:8787/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"gpt-4o-mini","stream":true,"messages":[{"role":"user","content":"hello"}]}'流式响应注入压缩统计:data: {"headroom_stats":{"original_tokens":...,"compressed_tokens":...,"savings":...}}
| Integration | Package | Description |
|---|---|---|
| OpenAI Go SDK | integrations/go-openai |
HTTP RoundTripper 透明压缩 |
| langchaingo | integrations/langchaingo |
Document Compressor |
| Ollama | integrations/ollama |
HTTP 中间件 |
// OpenAI Go SDK
import headroom "github.com/superops-team/headroom-go/integrations/go-openai"
client := headroom.WrapClient(http.DefaultClient, headroom.Config{Aggressiveness: 0.5})// langchaingo
import headroom "github.com/superops-team/headroom-go/integrations/langchaingo"
compressor := headroom.NewDocumentCompressor(headroom.Config{Aggressiveness: 0.5})
docs, _ := compressor.CompressDocuments(ctx, documents)headroom mcp serve # 4 tools: compress, retrieve, stats, read兼容 Claude Code、Codex、Cursor 等所有 MCP 客户端。
headroom wrap claude --apply # Claude Code
headroom wrap codex --apply # OpenAI Codex
headroom wrap copilot # GitHub Copilot CLI
headroom wrap generic # 通用配置docker pull ghcr.io/superops-team/headroom-go:latest
docker run -p 18787:18787 ghcr.io/superops-team/headroom-go:latest镜像 <10MB(FROM scratch),支持 linux/amd64 + linux/arm64。
# Helm
helm repo add headroom https://superops-team.github.io/headroom-go
helm install headroom headroom/headroom
# Sidecar
kubectl apply -f https://raw.githubusercontent.com/superops-team/headroom-go/main/integrations/k8s/sidecar.yaml[Service]
Environment=HEADROOM_API_KEY=sk-xxx
ExecStart=/usr/local/bin/headroom proxy --port 8787
Restart=alwaysGitHub Actions 自动化:ci.yml(build + test + lint + security)、release.yml(多平台二进制 + Docker 推送)、docs.yml、benchmark.yml。
Your App Headroom Go LLM API
────────── ────────────── ──────────
│ Tool outputs │──→ │ Auto-detect │──→ │ Compressed │
│ Logs │ │ content type │ │ messages │──→ OpenAI
│ RAG snippets │ │ Apply best │ │ (70% fewer │ Anthropic
│ Code diffs │ │ compressor │ │ tokens!) │ etc.
─────────────── ──────────────── ───────────────
Two paths: Legacy (fast, default) and Pipeline (policy-driven, budget-aware).
| Type | Detection | Strategy |
|---|---|---|
| JSON | {/[ + valid |
Remove nulls, fold arrays, truncate floats |
| Code | 3+ lines with keywords | Strip comments, fold long functions |
| Text | Default | Deduplicate lines, remove stopwords |
| Diff | @@ headers |
Collapse unchanged hunks |
| Log | ERROR/WARN/INFO patterns | Preserve FATAL/ERROR, fold repeated |
| Search | filename:line: format |
Collapse repeats, preserve grouping |
| Tabular | TSV/CSV/table | Header-preserving text fallback |
| Spreadsheet | Reserved | Cell-level (planned) |
| HTML | <!doctype>/<html> |
Strip comments/script/style |
| Unknown | Fallback | Text compression |
| Benchmark | Throughput |
|---|---|
| Content Detection (1MB) | 390 MB/s |
| Tokenizer (1MB) | 95 MB/s |
| Diff Compressor (5k lines) | 6.1M lines/s |
| Log Compressor (50k lines) | 2.3M lines/s |
| End-to-End (mixed) | 650 ops/s |
go test -bench=. -benchtime=1s ./...git clone https://github.com/superops-team/headroom-go.git
cd headroom-go
go test -race -count=1 ./... # All tests
go vet ./... # Static analysis
go test -cover ./... # Coverage (85.2%)
go build -o headroom ./cmd/headroom # BuildContributions welcome! See CONTRIBUTING.md.
- Fork → 2. Branch → 3. Commit → 4. Push → 5. PR
go test -race ./... && go vet ./...MIT — see LICENSE.
Built with ❤️ by superops-team · Pure Go · Zero deps · MCP Native · Docker <10MB