Open source Apache-2.0 Local-first Built on Pi

Same context. Different models. Better decisions.

Sidebet is an open-source harness around your AI coding agent. It keeps your rules and memory consistent across sessions, runs the same task on different models side by side in isolated Git worktrees, and shows you the real diffs, test results and token costs before you keep one. Over time, your own history becomes evidence about which model works for your projects.

Get started View on GitHub
sidebet compare "Refactor the search handler"
~/acme-shop sidebet compare
sidebet v0.1.0 acme-shop · main
2 attempts · 2 running · 0s ctrl+c cancel
task Make search case-insensitive and add a test for it
A anthropic/claude-sonnet-5-5
⠋ running · 0s
tokens0 ↑0 ↓0
turns0
tools0
costcost n/a
creating worktree
B openai/gpt-5.5
⠋ running · 0s
tokens0 ↑0 ↓0
turns0
tools0
costcost n/a
creating worktree
Compare run_7f3a9c21 completed
 
A anthropic/claude-sonnet-5-5 completed 2 files changed tests: passed (14 passed) 2m 58s 6 turns, 14 tools 48.2k tokens $0.21
B openai/gpt-5.5 completed 3 files changed tests: passed (14 passed) 3m 31s 9 turns, 21 tools 61.7k tokens cost n/a
B: 1 tool call blocked by policy
 
next sidebet review run_7f3a9c21 your working tree was not modified

Illustrative output. The numbers come from a demo run, not a benchmark, and the models are examples.

Local-first SQLite on disk. No account, no telemetry, no cloud.
Apache-2.0 Open source, with a patent grant. Fork it, ship it.
Built on Pi Uses Pi's public SDK and the logins you already have.
Your tree, untouched Comparisons run in Git worktrees from a snapshot.

The problem

One agent. One answer. No way to know if it was the best one.

If you use an AI coding agent every day, with Claude Code, Codex, Cursor, Pi or anything else, you have probably felt all three of these.

01

You re-explain your project every session.

Conventions, slow test suites, files that must not change: they live in your head and in whichever chat you had open. Every new session starts from zero.

Rules and memory are injected into every session. You can see exactly what the model received, and why each item was included or left out.

02

You pick a model on vibes.

Trying another model means redoing the task by hand and juggling branches. So you don't, and you never learn what you missed.

The same task runs on two or more models at once, each in its own Git worktree. Real diffs, real tests, real token counts, side by side.

03

Nobody keeps score.

Which model passes your tests more often? Which one do you actually prefer once you read the diff? Today that is a feeling.

Every run is recorded. sidebet odds shows pass rates, your preferences and medians per model, with sample sizes, never blended into a score.

What it is

Not another coding agent.

Pi runs the agent. Sidebet owns everything around it: context, memory, policy, isolation, review and history. Pi is never forked or patched; Sidebet uses its public SDK and the credentials it already holds.

Your repository AGENTS.md · rules · tests · Git history
Sidebet the harness: a CLI and a thin Pi extension
contextmemorypolicyisolationreviewhistory
Pi, unmodified the agent: sessions, tools, skills, your logins
Your model providers Anthropic · OpenAI · GitHub Copilot · xAI · Kimi · and whatever else Pi can log into

What Sidebet is not

Saying this plainly is the point.

  • Not a model, a proxy or an inference layer. Prompts go to the providers you already logged into through Pi, exactly as they would from Pi.
  • Not a fork of Pi. It uses Pi's public SDK and reads the credentials, default model, skills and AGENTS.md Pi already has. Nothing to re-enter.
  • Not a benchmark. It reports what happened on your tasks, in your repo, with your tests. No leaderboards, no "best model".
  • Not a cloud service. Everything in v1 runs on your machine with no account. See the roadmap for what may come later.
  • Not a chat app or an IDE. It is a CLI built on Pi's own terminal toolkit, so it looks and behaves like Pi.

How it works

Six commands, from a fresh repo to evidence.

Everything below is implemented and tested today. The output shown is illustrative; the models are examples.

Init in seconds. Nothing to re-enter.

Sidebet inspects an existing repository and reports what it found: language, package manager, test command, an existing AGENTS.md, and how many models Pi has credentials for. It writes exactly one small file, never overwrites it, and is idempotent.

~/acme-shopinit
sidebet init
Sidebet initialized

Project      acme-shop
Runtime      pi 1.1.0
Language     TypeScript
Package mgr  pnpm
Rules        AGENTS.md detected
Models       7 available (default anthropic/claude-sonnet-5-5)
Tests        pnpm test

Wrote .sidebet/config.json (commit it to share the project id)

You're ready.

  sidebet run "Explain this repository"
  sidebet compare "Improve the search handler"

Run with context that follows you.

Rules are always injected. Memory is advisory, searched per task, and never overrides a rule. --dry-run shows the whole plan, every context item and the reason it was included, without calling a model. --readonly removes write tools in code, not in the prompt, and verifies the tree afterwards.

~/acme-shoprun --dry-run
sidebet memory add "Search must stay case-insensitive; never run regexes on user input"
Saved mem_fa704c29 (project). Included in relevant future runs; explicit rules still take precedence.

sidebet run "Make search case-insensitive" --dry-run
Dry run: nothing will be executed, no model will be called, nothing is saved.

Task          Make search case-insensitive
Mode          run in your working tree
Models        anthropic/claude-sonnet-5-5
Tools         read, bash, edit, write
Policy        pol_3a11f4fdbe16

Why each item
  native   agents  AGENTS.md                      ~410t  - loaded by the runtime itself
  injected rules   .sidebet/rules/conventions.md  ~260t  - always injected
  injected memory  mem_fa704c29                   ~21t   - included: matched task

Policy decisions
  protected paths: .env, .env.*, **/*.pem, **/*.key, **/id_rsa*, .git/**
  approval required (bash): rm -rf, git push, git reset --hard, git clean, sudo, curl | sh

Estimated context: 691 tokens (281 injected by Sidebet, 410 loaded by the runtime, budget 12,000)

Compare, safely.

Two or more models run the same task concurrently, each in its own Git worktree created from a snapshot of your working tree, uncommitted changes included. Your files are never touched. Each attempt gets your test suite. Afterwards a metrics table puts them side by side; ◂ marks the better value for that one metric, and nothing is combined into a score.

~/acme-shopcompare
sidebet compare "Make search case-insensitive and add a test for it" --models sonnet,gpt5
Comparing anthropic/claude-sonnet-5-5 vs openai/gpt-5.5
Attempts start from a snapshot of your working tree (2 uncommitted changes included). Your files are untouched.
…

Metrics
A claude-sonnet-5-5B gpt-5.5
Statuscompletedcompleted
Total time2m 58s ◂3m 31s
Agent time2m 31s ◂3m 02s
Time to first token1.4s ◂2.1s
Time to first edit41s ◂1m 12s
Turns69
Tool calls14 (read 6, edit 3, bash 3, grep 1, write 1)21 (read 8, edit 5, bash 4, find 2, grep 1, write 1)
Tool errors / blocked0 / 00 / 1
Input tokens45.1k ◂57.9k
Output tokens3.1k ◂3.8k
Total tokens48.2k ◂61.7k
Cost$0.21not reported
Files changed23
Lines +/-+38 / -12+71 / -19
Testspassed (14 passed)passed (14 passed)
  ◂ marks the better value for that single metric; metrics are not combined into a score.
  Cost is what the runtime reported; subscription logins often report none, and nothing is estimated.

next  sidebet review run_7f3a9c21   your working tree was not modified

Review with real metrics. Apply with an undo.

Diffs open in a full-screen viewer with Pi's own rendering. You record which result you prefer and why. Applying previews the files, warns about your local edits, refuses on conflict, saves a backup ref first, changes only the working tree, and prints the command to undo it. Nothing is ever merged automatically.

~/acme-shopreview
sidebet review
run_7f3a9c21 · A claude-sonnet-5-5 vs B gpt-5.5

  ▸ View A diff
    View B diff
    Compare A and B
    Select preferred result
    Record review                       accepted or not, with a reason
    Apply a result to this repository
    Check out an attempt in a worktree
    Exit

Recorded rev_48d693c9 · preferred A · accepted · "smaller diff, kept the existing test structure" · your call ◉ matched

Apply A (anthropic/claude-sonnet-5-5) to /Users/you/acme-shop
     src/search.ts
     test/search.test.ts
Applies to the working tree only (no commit, index untouched). A backup ref of your current tree is saved first.
Apply these changes? [y/N] y
Applied 2 files.
  Attempt kept at:   refs/sidebet/runs/run_7f3a9c21/A
  To undo:           git checkout refs/sidebet/backups/run_7f3a9c21 -- "src/search.ts"   (for files that existed before)

Learn from your own history.

Per model: runs, test pass rate (objective), how often you preferred it (human), acceptance, median time, tokens and cost. Sample sizes are always shown. The head-to-head table only uses runs where both models attempted the same task. sidebet bet A records your prediction before you read the diffs; odds reports how often your calls matched, a measure of your judgement, not of the models.

~/acme-shopodds
sidebet odds
▗▟▀▙▖  sidebet v0.1.0  odds · acme-shop
▝▜▄▛▘  31 finished runs · 31 compare · 26 with a recorded preference

Model                        Runs  Tests passed (objective)  Preferred (human)      Accepted     Median time  Median tokens  Median cost
anthropic/claude-sonnet-5-5  31    ███████░ 27/31 (87%)      ██████░░ 19/26 (73%)  22/26 (85%)  3m 04s       52.1k          $0.24
openai/gpt-5.5               31    ██████░░ 24/31 (77%)      ██░░░░░░ 7/26 (27%)   18/26 (69%)  3m 48s       64.9k          n/a

Head-to-head (same task, both attempted)
  claude-sonnet-5-5 vs gpt-5.5: 31 runs; preferred 19-7; tests: 22 both passed, 5 only sonnet, 2 only gpt-5.5, 2 both failed (n=31)

Your calls  ◉◉◎◉◉◉◎◉◉◉◉◉  21/26 matched your later preference
  A call is a `sidebet bet` placed before reviewing; it measures your prediction, not the models.

• Test pass rate and preference are different evidence: tests are objective but only as good as your suite;
  preference is your judgement. They are shown separately and never blended.
• Cost is shown only where the runtime reported one; missing costs are not estimated.

See everything in a browser. Locally.

sidebet ui serves a read-only web view of the same records: every run, attempts side by side, metrics, per-file diffs, tool calls, test output, the exact context the model received, odds and memory. It binds to 127.0.0.1, has no write endpoints, and adds no dependencies.

~/acme-shopui
sidebet ui
▗▟▀▙▖  sidebet v0.1.0  ui · acme-shop
▝▜▄▛▘  http://127.0.0.1:4242/  ctrl+c to stop

Read-only. Serves this project's runs, attempts, diffs, context snapshots, odds and memory
from ~/.sidebet/sidebet.db to your browser. Nothing leaves your machine.

sidebet runs
ID            STATUS     KIND     MODELS                     WHEN     TASK
run_7f3a9c21  completed  compare  sonnet-5-5 vs gpt-5.5      2m ago   Make search case-insensitive and add a test for it
run_c41e08d2  completed  compare  sonnet-5-5 vs gpt-5.5      3h ago   Add input validation to the checkout endpoint
run_9be0a7f3  completed  run      sonnet-5-5                 1d ago   Explain the auth flow read-only
run_2d7c5e10  timeout    compare  sonnet-5-5 vs gpt-5.5      2d ago   Migrate the search index to the new schema

Everything around the agent

The harness, command by command.

Each of these exists today and is covered by tests that run without paid model calls.

sidebet init

Detects, never dictates

Git, package manager, test command, AGENTS.md, Pi models and skills. Writes one small config file. Idempotent.

sidebet run

One task, full receipts

A Pi-backed session with your rules and memory injected, tool activity streamed, usage and the exact change set recorded.

--readonly · --dry-run

Safe modes that are real

Read-only exposes only read, grep, find and ls, blocks writes in code and verifies the tree afterwards. Dry run prints the whole plan and calls nothing.

sidebet compare

Isolated, concurrent, measured

Several models, one task, each in its own worktree from a snapshot of your tree. Tests per attempt. A live dashboard while it runs.

sidebet review

Metrics, diffs, verdict, apply

Column-by-column metrics, full-screen diffs, a recorded review, and an apply that previews, backs up, refuses conflicts and prints the undo.

sidebet odds · bet

Your own evidence

Pass rates, preferences and medians per model, sample sizes always shown. Place a call before reviewing and learn how good your instincts are.

sidebet runs · rerun

History you can replay

Every run, attempt and diff. Rerun with the original context snapshot or the current one, stated explicitly. Inputs are reproduced, not outputs.

sidebet context

Transparent context

What the agent would receive now, or the stored snapshot of any past run: each item injected, native or omitted, its tokens, and why.

sidebet memory

Memory that knows its place

Project and global memory in SQLite with full-text search. Advisory: rules always win. Suggestions are shown, never saved without you.

sidebet policy check

Policy, enforced in code

Provider and model allow-lists, tool deny lists, protected paths and approval-required commands. The matrix says what is enforced, partial or unsupported.

sidebet ui · doctor

Local web view, honest diagnostics

A read-only browser view of runs, diffs, context, odds and memory on 127.0.0.1. Doctor explains what Sidebet sees and why something was refused.

pi install ./extensions/pi

Inside Pi too

/sidebet-compare, /sidebet-review and /sidebet-odds inside Pi, plus your rules and relevant memory added to plain Pi sessions.

Trust

Read this before trusting it.

Sidebet separates instructions to a model from enforcement in code. A model can ignore an instruction. It cannot ignore a hook that blocks the tool call before it runs. Here is which is which.

✓ Enforced in code

  • Model and provider policy is checked before a session starts. Denied models are never contacted.
  • Read-only mode: write-capable tools are not exposed, a tool-call hook blocks them anyway, user extensions are not loaded, and the tree is compared before and after.
  • Tool allow and deny lists, protected path patterns for file tools, and approval-required commands that pause for a human on a terminal and are denied otherwise.
  • Secret redaction for well-known shapes (private keys, cloud and API tokens, KEY=value lines) before anything is injected or stored.
  • Fail closed: if the policy hook cannot load inside Pi, the attempt does not run.

! Not enforced. Be aware.

  • A shell can read anything. Commands that literally name a protected path are blocked, best effort. policy.strict removes bash entirely.
  • Pi extensions and skills run with your permissions, outside Sidebet policy except in read-only mode.
  • Redaction is pattern-based. Novel secret formats will not be caught. Do not put secrets in rules or memory.
  • No sandbox. Agents run as you. Worktrees isolate Git state, not the filesystem or the network.
  • Local data is unencrypted. Run records, diffs and redacted context live in ~/.sidebet/sidebet.db. Prompts go to the providers you chose, exactly as with Pi.

The complete list is in docs/security.md. sidebet policy check prints the same matrix for your own configuration.

Roadmap

Where this is going.

Open core. The CLI, the core library and the Pi adapter are Apache-2.0 and stay that way. Everything in the first column exists; the rest is intent, not a promise.

Now · v0.1

Shipped and local

Everything on this page runs on your machine today.

  • init, run, compare, review, odds, bet, runs, rerun
  • Context inspection and persistent memory with full-text search
  • Policy: read-only, allow and deny lists, protected paths, approvals
  • Live dashboard, metrics table, full-screen diff viewer
  • Pi extension: /sidebet-compare, /sidebet-review, /sidebet-odds
  • Local read-only web view with sidebet ui
Later · Sidebet Cloud

Optional, for teams

Not built. The local CLI stays complete without it.

  • Sync of runs and evaluation history across machines
  • Shared team memory and rules
  • An organisation policy layer above your local policy
  • Hosted evaluation history and a team console
  • A second runtime behind the same AgentRuntime interface
  • Versioned protocol and SDK packages

Non-goals for now: accounts, cloud sync, a vector store, a plugin marketplace, automatic memory promotion, automatic merging.

Would Sidebet matter to your team?

Shared memory, team policy and hosted history are the cloud product we have not built. If you want it, say so. Interest decides what gets built next.

No list, no tracking. One email when there is something to show.

Get started

Your first comparison in about twenty minutes.

If Pi runs on your machine, Sidebet runs.

Sidebet reuses your Pi login and your AGENTS.md. There are no accounts and nothing to re-enter. The walkthrough below ends with you comparing two models on a real change and choosing which result to keep.

Node ≥ 22.19pnpmGitPi 1.1macOS · Linux
Install from source
git clone https://github.com/jamesray/sidebet.git
cd sidebet
pnpm install && pnpm build
cd apps/cli && pnpm link --global   # puts `sidebet` on your PATH
sidebet doctor
Your first comparison
cd your-project
sidebet init
sidebet models                                   # the models you are logged into in Pi
sidebet compare "Make search case-insensitive" --models <a>,<b>
sidebet review
Optional: inside Pi
pi install ./extensions/pi   # adds /sidebet, /sidebet-compare, /sidebet-review, /sidebet-odds

Stop guessing which model to trust.

Run the task on both. Read the diffs. Keep the receipts.