Background
Codebuff is an AI programming assistant, with its free tier Freebuff providing several models (Gemini, DeepSeek, GLM, Kimi, etc.) that can be called for free. Officially, only a CLI client is provided, and no public API is available.
Goal: Reverse engineer Freebuff’s free models to conform to the standard /v1/chat/completions (OpenAI protocol) and /v1/messages (Claude protocol), so that any compatible client (LobeChat, NextChat, Claude Code, Codex, etc.) can be used directly.
1. Protocol Overview: Four Key Endpoints
By capturing real CLI traffic, the full Freebuff backend call chain was mapped out. All requests include Authorization: Bearer <authToken>.
| Step | Endpoint | Purpose |
|---|---|---|
| 1 | POST /api/v1/freebuff/session |
Create/keep alive a free session, returns instanceId and expiresAt |
| 2 | POST /api/v1/agent-runs (action=START) |
Start an agent run, returns runId |
| 3 | POST /api/v1/chat/completions |
Actual chat request, injects codebuff_metadata into payload |
| 4 | POST /api/v1/agent-runs (action=FINISH) |
End the run, report steps/credits |
1. Session Creation
POST /api/v1/freebuff/session
Authorization: Bearer <token>
x-freebuff-model: deepseek/deepseek-v4-flash
Content-Type: application/json
{}
Response:
{
"status": "active",
"instanceId": "xxxx",
"model": "...",
"expiresAt": "2026-07-27T...",
"rateLimit": {...}
}
Note: this is a POST with {} empty body + x-freebuff-model header, not GET. Sessions might enter a queue (status: "queued", with position/queueDepth), polling is needed.
2. Starting a Run
POST /api/v1/agent-runs
Authorization: Bearer <token>
{
"action": "START",
"agentId": "base2-free",
"ancestorRunIds": []
}
Response:
{"runId": "2b56444d-..."}
3. Chat Request (key point: injecting metadata)
The chat payload is basically in OpenAI format, but you must inject codebuff_metadata—all four fields are required:
{
"model": "google/gemini-2.5-flash-lite",
"messages": [...],
"stream": false,
"codebuff_metadata": {
"run_id": "<runId from step 2>",
"cost_mode": "free",
"client_id": "<random 13-digit hex>",
"freebuff_instance_id": "<instanceId from step 1>"
}
}
2. Pitfall Log: Three Key Traps
Trap 1: free_mode_invalid_agent_hierarchy — Run must be attached to the root session
At first, a run was STARTed per agent, but all chat attempts were rejected:
{"error":"free_mode_invalid_agent_hierarchy","message":"Free mode subagents must run under an active freebuff session root."}
Root cause: Freebuff enforces a tree structure with session as root:
Session (instanceId)
└── Root run: base2-free ← This must be created first, ancestorRunIds: []
├── Sub run: file-picker ← ancestorRunIds: [root runId]
└── Sub run: code-reviewer-* ← ancestorRunIds: [root runId]
A sub-run’s ancestorRunIds can only include the root run’s id. The original mistake was putting all sibling run IDs in ancestor list, which was immediately flagged as invalid.
Fix: For each token, keep one and only one root run (base2-free), sub-runs are lazy-created and only reference the root as ancestor.
Trap 2: 400 Invalid request body — null vs
For the root run with no ancestor, what should ancestorRunIds be? In Go, var ancestors []string is nil, serialized by json.Marshal as null:
{"ancestorRunIds": null}
Backend immediately returns 400:
{"error":"Invalid request body","details":{"ancestorRunIds":{"_errors":["Invalid input: expected array, received null"]}}}
Fix: Initialize as an empty slice ancestors := []string{} so it serializes as [] not null. Classic strong-typed vs weak-typed schema bug.
Trap 3: free_mode_invalid_agent_model — Free models now severely restricted
According to the official open source repo’s free-agents.ts, the free tier should support many models. Registered them all, but all but Gemini were rejected:
{"error":"free_mode_invalid_agent_model","message":"Free mode is only available for specific agent and model combinations."}
Testing found (as of 2026-07):
| Model | Result |
|---|---|
google/gemini-2.5-flash-lite |
|
google/gemini-3.1-flash-lite-preview |
|
| deepseek-v4-pro/flash, minimax-m3, glm-v5.2, kimi-k2-thinking, mimo-v2.5, hy3, laguna-s-2-1 |
Lesson: free-agents.ts source ≠ backend reality. Free tier is locked down, only use empirically tested models, not just what docs say.
3. CLI Detection: Why Cloudflare Worker Doesn’t Work
Initially tried to run the proxy as a Cloudflare Worker (no infra, global distribution). But all requests from Worker were rejected, no matter the disguise:
{"error":"free_mode_cli_required"}
All kinds of spoofing failed:
User-Agent: Freebuff-CLI/0.0.105(exact same as official CLI)- Adding
Origin/Referer/Host/Accept-Encodingetc. matching CLI/browser - Removing all headers revealing Worker identity
- Handling gzip responses
Conclusion: Detection is not at HTTP header level, but TLS fingerprint layer (Client Hello / JA3). Cloudflare Worker fetch uses Cloudflare’s own TLS stack, fingerprint differs completely from Node.js’s undici/OpenSSL and cannot be spoofed at header level.
Solution: The local Go binary uses Go’s native TLS stack, and its fingerprint isn’t blocked—request goes through. So the final solution is a local proxy service rather than serverless.
4. Final Architecture
┌─────────────┐ OpenAI/Claude Protocol ┌──────────────────┐ Freebuff Private Protocol ┌────────────┐
│ Any Client │ ────────────────────────▶ │ Freebuff2API │ ────────────────────────────▶ │ Codebuff │
│ (LobeChat…) │ ◀──────────────────────── │ (Go, local :8080)│ ◀────────────────────────── │ Backend │
└─────────────┘ └──────────────────┘ └────────────┘
│
┌───────────────────────┼───────────────────────┐
│ ModelRegistry │ RunManager │ SessionPool
│ Hardcoded + Dynamic │ One root run/token │ Session keepalive/queue
│ agent→model mapping │ Lazy-create sub-runs │ Auto-refresh on expiry
└───────────────────────┴───────────────────────┘
Key Design Points
1. Model Registry: Hardcoded base, remote supplement
Upstream free-agents.ts now uses FREEBUFF_*_MODEL_ID constants, so regex can’t directly parse literals anymore. Strategy: hardcode an empirically tested agent→model mapping as the authoritative base; merge in remote results incrementally, ensuring only additions.
2. Run Lifecycle Management
- Only warm up 1 root run on start (not all 28 agents to avoid pointless START/FINISH storms).
- Sub-runs are created only when first needed, ancestor always points to root.
- On root run expiry and rotation, sub-runs are rotated to new root.
- Release lease at request end; FINISH old root runs if emptied out.
3. Session Management
- Cache one session per token, auto-refresh 5 seconds before expiry.
- If queued, return
Retry-Afterto client (don’t just wait blindly). - On 401, put token in 30-min cooldown to prevent repeated failures.
4. Dual Protocol Endpoints
/v1/chat/completions: Standard OpenAI, supports both streaming and non-streaming./v1/messages: Claude protocol, does both request/response conversion./v1/models: Returns empirical list of available models.
5. Empirical Testing
GET /v1/models
→ gemini-2.5-flash-lite, gemini-3.1-flash-lite-preview
POST /v1/chat/completions (gemini-2.5-flash-lite, "Say OK")
→ {"choices":[{"message":{"content":"Alright"}}], "usage":{...}}
POST /v1/chat/completions (stream: true, "17*23?")
→ data: {...391...} data: [DONE] ← Streaming works fine
POST /v1/messages (Claude protocol)
→ {"content":[{"text":"All right.","type":"text"}], "role":"assistant", ...}
6. Takeaways
- Packet capture > docs. Just because
free-agents.tssays something doesn’t mean backend will allow it. Always trust what you test yourself. - Private protocol constraints are often in hidden fields. All four fields for
codebuff_metadataand hierarchy inancestorRunIds—miss any, get 400. null≠[]. Go’s nil slice serializes to null and is rejected by backend schema—use empty slice instead.- TLS fingerprint is a hard barrier. Serverless (Worker) TLS stack is identified, only native local TLS passes—deployment model is determined by this.
- Model hierarchy explicitly. Code your “session → root run → sub-run” tree, with a dedicated field for root run; don’t just flatten in a map by convention. This prevents a whole class of bugs.
Project URL
- Freebuff2API:
- Single-file Go binary; fill in
config.jsonwith authToken and you’re good to go - Get token at: https://freebuff.llm.pm (shown directly after login)
Disclaimer: This project is for educational and research purposes only, not officially affiliated with Codebuff/Freebuff. Please follow upstream service terms and keep call frequency reasonable.