Unlimited Free Google/Gemini-3.1-Flash-Lite-Preview (Reverse Engineering Special Thread)

Background

Codebuff is an AI programming assistant, with its free tier Freebuff providing several models (Gemini, DeepSeek, GLM, Kimi, etc.) that can be called for free. Officially, only a CLI client is provided, and no public API is available.

Goal: Reverse engineer Freebuff’s free models to conform to the standard /v1/chat/completions (OpenAI protocol) and /v1/messages (Claude protocol), so that any compatible client (LobeChat, NextChat, Claude Code, Codex, etc.) can be used directly.


1. Protocol Overview: Four Key Endpoints

By capturing real CLI traffic, the full Freebuff backend call chain was mapped out. All requests include Authorization: Bearer <authToken>.

Step Endpoint Purpose
1 POST /api/v1/freebuff/session Create/keep alive a free session, returns instanceId and expiresAt
2 POST /api/v1/agent-runs (action=START) Start an agent run, returns runId
3 POST /api/v1/chat/completions Actual chat request, injects codebuff_metadata into payload
4 POST /api/v1/agent-runs (action=FINISH) End the run, report steps/credits

1. Session Creation

POST /api/v1/freebuff/session
Authorization: Bearer <token>
x-freebuff-model: deepseek/deepseek-v4-flash
Content-Type: application/json

{}

Response:

{
  "status": "active",
  "instanceId": "xxxx",
  "model": "...",
  "expiresAt": "2026-07-27T...",
  "rateLimit": {...}
}

Note: this is a POST with {} empty body + x-freebuff-model header, not GET. Sessions might enter a queue (status: "queued", with position/queueDepth), polling is needed.

2. Starting a Run

POST /api/v1/agent-runs
Authorization: Bearer <token>

{
  "action": "START",
  "agentId": "base2-free",
  "ancestorRunIds": []
}

Response:

{"runId": "2b56444d-..."}

3. Chat Request (key point: injecting metadata)

The chat payload is basically in OpenAI format, but you must inject codebuff_metadata—all four fields are required:

{
  "model": "google/gemini-2.5-flash-lite",
  "messages": [...],
  "stream": false,
  "codebuff_metadata": {
    "run_id": "<runId from step 2>",
    "cost_mode": "free",
    "client_id": "<random 13-digit hex>",
    "freebuff_instance_id": "<instanceId from step 1>"
  }
}

2. Pitfall Log: Three Key Traps

Trap 1: free_mode_invalid_agent_hierarchy — Run must be attached to the root session

At first, a run was STARTed per agent, but all chat attempts were rejected:

{"error":"free_mode_invalid_agent_hierarchy","message":"Free mode subagents must run under an active freebuff session root."}

Root cause: Freebuff enforces a tree structure with session as root:

Session (instanceId)
└── Root run: base2-free           ← This must be created first, ancestorRunIds: []
    ├── Sub run: file-picker       ← ancestorRunIds: [root runId]
    └── Sub run: code-reviewer-*  ← ancestorRunIds: [root runId]

A sub-run’s ancestorRunIds can only include the root run’s id. The original mistake was putting all sibling run IDs in ancestor list, which was immediately flagged as invalid.

Fix: For each token, keep one and only one root run (base2-free), sub-runs are lazy-created and only reference the root as ancestor.

Trap 2: 400 Invalid request body — null vs

For the root run with no ancestor, what should ancestorRunIds be? In Go, var ancestors []string is nil, serialized by json.Marshal as null:

{"ancestorRunIds": null}

Backend immediately returns 400:

{"error":"Invalid request body","details":{"ancestorRunIds":{"_errors":["Invalid input: expected array, received null"]}}}

Fix: Initialize as an empty slice ancestors := []string{} so it serializes as [] not null. Classic strong-typed vs weak-typed schema bug.

Trap 3: free_mode_invalid_agent_model — Free models now severely restricted

According to the official open source repo’s free-agents.ts, the free tier should support many models. Registered them all, but all but Gemini were rejected:

{"error":"free_mode_invalid_agent_model","message":"Free mode is only available for specific agent and model combinations."}

Testing found (as of 2026-07):

Model Result
google/gemini-2.5-flash-lite :white_check_mark:
google/gemini-3.1-flash-lite-preview :white_check_mark:
deepseek-v4-pro/flash, minimax-m3, glm-v5.2, kimi-k2-thinking, mimo-v2.5, hy3, laguna-s-2-1 :cross_mark: All rejected

Lesson: free-agents.ts source ≠ backend reality. Free tier is locked down, only use empirically tested models, not just what docs say.


3. CLI Detection: Why Cloudflare Worker Doesn’t Work

Initially tried to run the proxy as a Cloudflare Worker (no infra, global distribution). But all requests from Worker were rejected, no matter the disguise:

{"error":"free_mode_cli_required"}

All kinds of spoofing failed:

  • User-Agent: Freebuff-CLI/0.0.105 (exact same as official CLI)
  • Adding Origin / Referer / Host / Accept-Encoding etc. matching CLI/browser
  • Removing all headers revealing Worker identity
  • Handling gzip responses

Conclusion: Detection is not at HTTP header level, but TLS fingerprint layer (Client Hello / JA3). Cloudflare Worker fetch uses Cloudflare’s own TLS stack, fingerprint differs completely from Node.js’s undici/OpenSSL and cannot be spoofed at header level.

Solution: The local Go binary uses Go’s native TLS stack, and its fingerprint isn’t blocked—request goes through. So the final solution is a local proxy service rather than serverless.


4. Final Architecture

┌─────────────┐   OpenAI/Claude Protocol   ┌──────────────────┐   Freebuff Private Protocol   ┌────────────┐
│ Any Client  │ ────────────────────────▶ │  Freebuff2API    │ ────────────────────────────▶ │ Codebuff   │
│ (LobeChat…) │ ◀──────────────────────── │  (Go, local :8080)│ ◀────────────────────────── │ Backend    │
└─────────────┘                          └──────────────────┘                            └────────────┘
                                             │
                   ┌───────────────────────┼───────────────────────┐
                   │  ModelRegistry        │  RunManager           │  SessionPool
                   │  Hardcoded + Dynamic │  One root run/token   │  Session keepalive/queue
                   │  agent→model mapping  │  Lazy-create sub-runs │  Auto-refresh on expiry
                   └───────────────────────┴───────────────────────┘

Key Design Points

1. Model Registry: Hardcoded base, remote supplement

Upstream free-agents.ts now uses FREEBUFF_*_MODEL_ID constants, so regex can’t directly parse literals anymore. Strategy: hardcode an empirically tested agent→model mapping as the authoritative base; merge in remote results incrementally, ensuring only additions.

2. Run Lifecycle Management

  • Only warm up 1 root run on start (not all 28 agents to avoid pointless START/FINISH storms).
  • Sub-runs are created only when first needed, ancestor always points to root.
  • On root run expiry and rotation, sub-runs are rotated to new root.
  • Release lease at request end; FINISH old root runs if emptied out.

3. Session Management

  • Cache one session per token, auto-refresh 5 seconds before expiry.
  • If queued, return Retry-After to client (don’t just wait blindly).
  • On 401, put token in 30-min cooldown to prevent repeated failures.

4. Dual Protocol Endpoints

  • /v1/chat/completions: Standard OpenAI, supports both streaming and non-streaming.
  • /v1/messages: Claude protocol, does both request/response conversion.
  • /v1/models: Returns empirical list of available models.

5. Empirical Testing

GET /v1/models
→ gemini-2.5-flash-lite, gemini-3.1-flash-lite-preview

POST /v1/chat/completions  (gemini-2.5-flash-lite, "Say OK")
→ {"choices":[{"message":{"content":"Alright"}}], "usage":{...}}

POST /v1/chat/completions  (stream: true, "17*23?")
→ data: {...391...}  data: [DONE]   ← Streaming works fine

POST /v1/messages  (Claude protocol)
→ {"content":[{"text":"All right.","type":"text"}], "role":"assistant", ...}

6. Takeaways

  1. Packet capture > docs. Just because free-agents.ts says something doesn’t mean backend will allow it. Always trust what you test yourself.
  2. Private protocol constraints are often in hidden fields. All four fields for codebuff_metadata and hierarchy in ancestorRunIds—miss any, get 400.
  3. null ≠ []. Go’s nil slice serializes to null and is rejected by backend schema—use empty slice instead.
  4. TLS fingerprint is a hard barrier. Serverless (Worker) TLS stack is identified, only native local TLS passes—deployment model is determined by this.
  5. Model hierarchy explicitly. Code your “session → root run → sub-run” tree, with a dedicated field for root run; don’t just flatten in a map by convention. This prevents a whole class of bugs.

Project URL

  • Freebuff2API:
  • Single-file Go binary; fill in config.json with authToken and you’re good to go
  • Get token at: https://freebuff.llm.pm (shown directly after login)

Disclaimer: This project is for educational and research purposes only, not officially affiliated with Codebuff/Freebuff. Please follow upstream service terms and keep call frequency reasonable.

37 Likes

If that’s true, that would be awesome.

Looks good, you have my support!

But what can Hakimi even do?
(Are there really people who actually use this thing to write code? :xhj27: )

1 Like

Also, Cloudflare Worker might not work, but why not try edgeone? It also lets you deploy directly, but it’s a complete Node.js environment. If they block IPs, just set up an HTTP proxy to get around it.

If you want to unlock premium models, you need to use either residential proxy or vercel relay in supported country. I haven’t try cloudflare relay. Lots public vpn and cloudflare warp already blocked, they especially using third party service just to harden their detection.

1 Like

So impressive, I really envy those who work on reverse proxies and cracking.

A true mastermind.

Is my frontend skill really impressive?

2 Likes

That’s amazing—I don’t really understand it, but it sure looks impressive.

The information is already enough for getting started. It would be more useful to add common error messages related to “Background: Codebuff is an AI coding tool.” I’ll do a small test on my end before making a decision.

1 Like

Can it be run in batches?

Very detailed, worth learning from.

“Sounds impressive, even though I don’t really understand it.”

3 Likes

Study seriously!

Unlimited and still free?

2 Likes

Very detailed, I’ll study it.

Learn a bit

必须在部署机的浏览器访问网页获取token?

I’ve messed with this for quite a while, but it only supports login via Google and GitHub—mass registration isn’t possible, which is pretty useless. On top of that, the latency is quite high, so I just put it aside to collect dust.
image