🔬 Volunteer Compute
50M Idle AI Agents, One Public Queue: The Token Commons Is Now Open Source
The framework behind Free Startup Idea #151 is now public code: a task schema, a stdlib-only verifier, three example micro-tasks, and a reference queue site. Verification costs about half a percent of the human alternative.
Fifty million people pay for AI agents they barely use. Yesterday that observation was Free Startup Idea #151, The Token Commons: a public queue of bounded, verifiable scientific micro-tasks that any subscriber could point their own agent at, no provider permission required. Today the idea has a second artifact, because the framework is now open source.
Inside github.com/rayhe/token-commons is everything needed to launch a queue in a weekend: a machine-readable task schema, a stdlib-only verifier script, three example scientific micro-tasks with real verification, and a reference static queue site. Fork it and the commons multiplies, which is why it is open source. Volunteer computing never worked as a single project; SETI@home, Folding@home, and World Community Grid worked as a pattern anyone could join, and the ones that survived did so by making the join path trivial. This one works as a pattern anyone can fork.
What is actually in the repo
At the center sits schema/task.json, a JSON Schema for a task manifest. Every task declares the same contract: inputs (inline or URL plus sha256), natural-language instructions an agent can follow self-contained, an expected output format with typed fields, and a verification rule with an acceptance criterion. Four verification methods are supported, and each shifts trust from the solver to the method: deterministic tasks carry a precomputed sha256 of the correct answer; replication tasks require two independent submissions to agree byte-for-byte after canonicalization; spot-check tasks keep a gold table for sampled comparison; curator-review tasks name the reviewer. Any task that cannot be verified is refused by the schema, which is the entire point: tasks that cannot be verified do not ship.
verifier/verify.py is standard library only, so anyone can run it, and covers the whole claim flow in three subcommands: schema validates a manifest, check compares a submission's hash against the expected answer, and replicate compares two independent submissions. Canonicalization is fixed and documented: sort keys, compact separators, sha256. No ambiguity means no arguments about whose hash is right. The framework ships under MIT; task results default to CC0 so the commons' corpus cannot be enclosed later.
Three example tasks ship with the repo, each exercising a different verification method, each honest about what an agent can actually do. Task 001 is a catalog cross-match: 8 sky positions against a 14-entry reference catalog, find every pair within 2 arcseconds, with deliberate near-misses from 2.3 to 295 arcseconds to punish sloppy thresholds. For this task the curator precomputed the answer, and the manifest carries its hash. Task 002 extracts sample sizes, study designs, and effect estimates from three short synthetic abstracts, the canonical triage chore, verified by spot-check against a gold table; the abstracts are synthetic precisely so the gold table is undisputed, which is the only way a spot-check can be fair. Task 003 curates method, resolution, deposition date, and title for eight real PDB structures from RCSB, verified by replication: two agents must produce byte-identical output or the task returns to open.
The verification math nobody ran
Every critique of this design lands on the same point: verification is the whole game, and it is genuinely hard. That is true, so price it. Take the canonical queue chore, extracting sample sizes and effect estimates from 40 papers for a public-health review, and run both sides of the ledger with the inputs shown.
A graduate student takes roughly 15 minutes per paper to skim, extract, and tabulate: 10 hours total. At a loaded cost of about $35 an hour, that is $350. An agent processes roughly 40 papers at 6,000 tokens each, 240,000 input tokens, and emits an 8,000-token table. At $3 per million input tokens and $15 per million output tokens, the run costs about $0.84. Replication, the method that makes strangers' agents trustworthy, doubles it to about $1.70, and that number already includes the redundancy, because two agents doing the same work is the whole mechanism. Verification-inclusive agent cost lands near 0.5 percent of the human alternative, about four cents per verified record.
| Approach | Cost for 40 papers | Share of human cost |
|---|---|---|
| Grad student, 10 hrs at $35/hr | $350.00 | 100% |
| Single agent run | $0.84 | 0.24% |
| Agent run plus replication | $1.70 | 0.5% |
| Replication at $10/1M blended pricing | ~$5.00 | ~1.4% |
An honest caveat first: the human side assumes the chore is worth doing at all; nobody pays $350 to extract metadata from papers nobody will read. But that is exactly the demand profile the queue targets. Systematic literature extraction, docking triage, dataset labeling: data-rich and reviewer-poor, chores that die in backlogs because the human price is real and the budget is not. Beating a funded lab was never the bar. It needs to beat the backlog.
Limitations
This is a reference implementation, not a production system, and the gap between a skeleton that demonstrates the trust model and a queue that survives adversarial traffic is real. Claims ride on GitHub Issues in the skeleton; a live queue needs rate limiting per submitter, task locking while a claim verifies, and duplicate detection on result hashes, all documented as table stakes in the design but not built here. Small by construction, the three example tasks run tiny; the jump from 8 sky positions to 8 million is an engineering project the repo does not pretend to solve. Replication catches disagreement, not shared hallucination. Two agents trained on the same data can agree on the same wrong answer, which is why deterministic tasks with precomputed answers are the only method that fully sidesteps the problem, and why curators must keep writing them.
Modeled, not measured: the economics above show every input so readers can substitute their own. Token prices move, human loaded costs vary by institution, and the 15-minutes-per-paper estimate is a midpoint, not a law. Treat the 0.5 percent as an order of magnitude, not a quote; sensitivity checks keep it under 2 percent even at triple the token pricing. One more boundary: the design assumes human-paced, human-initiated solving. Automating task submission at scale would trip provider rate limits or terms of service, and a queue that depends on ban evasion is not a commons.
The strongest case against
Stated at full strength: this is Mechanical Turk with extra steps, and the steps are worse. Agents hallucinate, verification costs more than the volunteered compute is worth once you count curator time, and without a provider's blessing there is no quality control, just vibes. Volunteer computing's history ends in hibernation anyway: SETI@home stopped because the science ran out, not because the volunteers did. And the paid decentralized compute markets, Gensyn and Prime Intellect, are building the real version with prices instead of petitions. A volunteer agent queue is the charity gift shop next to their exchange.
Concede the verification burden, then refuse the conclusion.
Folding@home worked because work units were verifiable by construction, and the queue copies that constraint exactly: tasks that cannot be verified do not ship. But the Mechanical Turk comparison misses the price difference the table above makes concrete: the labor is already paid for by subscriptions, so the queue's only cost is verification, and verifying a bounded task is cheap next to the researcher's alternative, which is doing it by hand or not at all. SETI@home's diminishing returns were a data problem, the sky had been scanned; agent-shaped research is data-rich and reviewer-poor, the opposite profile, with demand that grows as the literature does. As for the decentralized markets, they are training markets with their own clients. Nobody has built the volunteer inference-task queue, because the only fleet of idle general-purpose agents is the subscriber base, and until today nobody had published the schema. Now it is public code.
What you can do
If you hold a paid AI subscription: no toggle needed and no permission required. Fork the repo, open the queue site, pick a task, and tell your agent to solve it: your agent, your account, your normal usage. The leaderboard is empty, and the first verified claim is still unclaimed.
If you run a lab: write three task manifests from the schema and open a PR. Your backlog of chores your grad students hate is the demand side, and the bar for every task is three words: bounded, verifiable, checkable.
If you are a builder: the verification tooling is the most valuable contribution you can make, because the schema is the standard and best tooling wins. Result-hash claiming, replication checks, rate limiting, duplicate detection: build the trust machinery.
If you are a curator: the scarcest resource in the design is judgment. Approve tasks against published criteria and reject vaporware in public with reasons, because that judgment function is the moat.
The Bottom Line
Every era finds its idle resource and builds a one-click way to put it to work. Screensavers in the 1990s, GPUs and office PCs in the 2020s. The idle resource of 2026 is stranger: not hardware at all, but millions of rented general-purpose agents sitting between their owners' questions, already paid for, already warmed up. Yesterday that idle pool was a startup idea. Today the schema, the verifier, and the queue are public code, and verification costs about half a percent of the human alternative. Post the queue, publish the schema, and let the agents solve.
Sources
- Free Startup Idea #151: The Token Commons, Live in the Future, September 28, 2026 (the origin of this article; published as defensive prior art)
- rayhe/token-commons, GitHub (the open-source framework: task schema, verifier, example tasks, reference queue site)
- OpenAI's Nick Turley says ChatGPT crossed 900M weekly users and 50M paying subscribers, Reuters, February 27, 2026
- Zooniverse passes 1 billion classifications, NASA Science, July 2026
- Folding@home passes 2.4 exaFLOPS, TechSpot, April 2020
- RCSB Protein Data Bank (data source for example task 003)