Caveman logo

Caveman

★★★★ 4.2/5
Visit site
Category
Coding
Pricing
Freemium

Quick Verdict

Caveman is not one terminal program. It is a local, open-source family built around reducing agent verbosity and token use. The basic Caveman skill tells Claude Code, Codex, Cursor, Windsurf, Cline, Copilot, and other agents to respond more tersely. Caveman Code is a separately installed full terminal coding agent with tool-output controls, read-only planning, autonomous goals, subagents, model selection, and sessions. Cavemem, Cavekit, the browser extension, SDKs, and MCP surfaces are separate components. The site also presents Gateway and Cloud work, but the hosted engine and control plane remain in development or waitlist status and should not be treated as equivalent to available MIT local tools.

Benchmark scope is the deciding issue. The Caveman skill’s “65% average reduction” covers output tokens only across ten prompts. It does not reduce input, context, files, or reasoning tokens, and its rules add about 1,000 to 1,500 input tokens per turn, making short replies potentially net-negative. Caveman Code’s “about 2x” comes from a separate 25-task MicroBench: with gpt-5.5 and xhigh reasoning, Caveman used 524k fresh tokens and passed 14/25, while Codex used 1,010k and passed 15/25, or about 1.93x fewer fresh tokens with one fewer passing task. This is a dated model, task set, and accounting method, not a universal performance warranty. Caveman is not marked as recommended; run an A/B against provider billing for your own work.

Best For

  • Agent users who regularly receive long explanations and will measure token use rather than trust a headline.
  • Developers seeking an inspectable MIT terminal agent with planning, subagents, and local sessions.
  • API users billed by token whose workloads contain long responses or large tool outputs.
  • Agent engineers studying prompt, tool-output, repetition, and memory compression.
  • Not ideal for per-request billing, already-terse replies, unreviewed remote installation, or buyers requiring a mature Cloud SLA.

Key Features

  • Caveman skill: lite, full, ultra, and wenyan styles shorten prose while preserving code, commands, paths, and error text.
  • Caveman Code: A separate terminal coding agent for interactive, print, tool-using, model, and session workflows.
  • Tool-output controls: Budgets read, grep, and shell output, removes ANSI noise, collapses whitespace, and avoids repeated full-file reads.
  • Planning gate: Plan mode exposes read-only tools before /act enters implementation.
  • Autonomous goals: Iteration, cost, no-progress, interruption, checkpoint, and ledger controls bound longer tasks.
  • Subagents and model roles: Worktree-isolated subagents and separate architect/editor models support parallel or lower-cost execution.
  • Cavemem and related components: Local memory, spec-driven development, browser mode, SDKs, and MCP can be installed and reviewed separately.
  • Multiple providers: API keys, compatible custom endpoints, and several OAuth paths are supported; every upstream retains its own terms and billing.

Use Cases

If the only requirement is shorter answers from an existing agent, test the skill before installing Caveman Code. Select five to ten normal tasks and record provider-billed input, cache, output, reasoning, requests, success rate, and elapsed time with and without the skill. Looking only at output hides rule injection, retries, tool calls, and context caching. Per-request products such as premium-request billing may charge the same amount for a shorter response.

Evaluate Caveman Code separately when a complete terminal agent is desired. It can read repositories, run commands, edit files, and send context to the selected model provider. Start in a test repository with non-production credentials, narrow tool permissions, inspect commands and diffs, and retain branches, tests, and review. Goal loops and subagents amplify both errors and spending, so set iteration, cost, directory, and network limits. Compare Claude Code, Codex, OpenCode, and Gemini CLI on the same tasks.

Pricing

The Caveman skill, Caveman Code, and several local components are MIT licensed and free to modify. The canonical 免费增值 classification reflects the same brand’s hosted Gateway, Engine, or Cloud path, which remains in development or waitlist status. A waitlist is not a purchasable production plan, and product demos do not establish a current SLA, region, quota, or data commitment. If hosted services launch, assess the live dashboard, agreement, privacy policy, and data-processing terms.

Real cost comes from model APIs, subscriptions, tool execution, storage, hardware, and review. Caveman Code can authenticate through existing subscriptions in some cases, but technical login support is not proof that every upstream authorizes every third-party client use. Check provider terms, quotas, and account risk. Token reductions also do not map to one fixed dollar saving because input, cache, reasoning, output, and requests can have different rates.

Pros

  • Separates a lightweight skill from a complete terminal agent, allowing minimal adoption.
  • MIT source makes prompts, scripts, runtime, and benchmark artifacts inspectable.
  • Benchmark documentation publishes scope and failure conditions rather than only repeating “65%.”
  • Caveman Code controls replies, tool output, and repeated reads, covering more than writing style alone.
  • Planning, checkpoints, subagents, MCP, skills, and multiple providers form a broad local toolkit.
  • Official documentation admits that terse work, per-request billing, and some agent accounting can be net-negative.

Cons

  • Skill rules add input every turn and may cost more for naturally short tasks.
  • The 65% output figure and 1.93x MicroBench fresh-token figure measure different products and cannot be combined.
  • The skill and Caveman Code are separate products under one umbrella, which can confuse feature attribution.
  • Remote install scripts, hooks, tool permissions, and OAuth increase local supply-chain and credential risk.
  • Local execution does not mean code stays local when a hosted model provider is selected.
  • Cloud remains in development or waitlist status, without a mature price, region, enterprise control set, or SLA to assess.

Alternatives

ToolBest forMain advantageWatch for
Claude CodeUsers prioritizing a mature terminal coding workflowFirst-party Anthropic agent experienceClosed service with separate billing and data terms
CodexOpenAI ecosystem and task executionIntegrated coding-agent and cloud workflowOne benchmark comparison is not an overall product ranking
OpenCodeOpen-client users seeking provider choiceOpen and extensible with flexible modelsDoes not provide the identical Caveman compression stack by default
Gemini CLIGoogle development workflowsFirst-party Gemini access and open terminal clientGoogle determines account, quota, and data processing
AiderGit-centered, multi-model code editingMature repository map and editing formatsDifferent interaction, autonomous-loop, and compression design

FAQ

Are Caveman and Caveman Code the same program?

No. Caveman is a terse-output skill installed into many agents. Caveman Code is a separate terminal coding agent. They share branding and a token-efficiency thesis but have different installation, capability, and data boundaries.

What exactly does “65% fewer tokens” mean?

It is the average output-token reduction for the Caveman skill across ten prompts. It excludes input, context, files, and reasoning, while the skill itself adds roughly 1,000 to 1,500 input tokens per turn.

Does “about half the tokens of Codex” apply to every task?

No. It came from 25 specific tasks under one model and reasoning setting. Caveman passed 14 and Codex passed 15. Reproduce the comparison on your repository and provider bill.

Is Caveman completely free?

The local MIT components are free software. Model APIs, subscriptions, hardware, and review still cost money. Hosted Cloud work is not covered by an assumption that MIT software guarantees a free hosted service.

Can locally running Caveman send code to a third party?

Yes. The skill is processed by its host, and Caveman Code calls the model provider selected by the user. Unless a verified local model is used, assume relevant prompts and code context can reach an upstream provider.

Should I run the one-line installation script directly?

That is a user risk decision. High-trust environments should download, inspect, and pin the script or release first. Confirm which agent directories, settings, and hooks it modifies and how removal works.

Bottom Line

Caveman’s useful idea is broader than one “65%” claim: terse replies, tool budgets, read deduplication, planning, memory, and a full terminal agent are exposed as inspectable components, with official documentation that acknowledges losing workloads. A proper review separates the skill, Caveman Code, other local tools, and the still-developing Cloud, then uses provider-billed A/B data. Long prose and large tool output may benefit materially. Short Q&A, per-request pricing, or repeated rule injection may cost more. Adopt the smallest component first, pin versions, use a test repository, and grant minimal permissions.

Last updated: July 21, 2026

Related tools