Methodology
How the survey runs, what an entry page is allowed to say, and where the numbers stop.
The Oblique-leaved Begonia, from The Temple of Flora, Robert John Thornton (ed.); painted by Philip Reinagle and others, 1799–1807. Amgueddfa Cymru — National Museum Wales (image 132043). Public domain. Credits
Part I
The survey
How one campaign runs, and what keeps the map from going stale.
The American Aloe, from The Temple of Flora, Robert John Thornton (ed.); painted by Philip Reinagle and others, 1799–1807. Amgueddfa Cymru — National Museum Wales (image 132006). Public domain. Credits
One campaign, start to finish
A campaign is one task, run once, on the record. The question and scope are frozen before the search starts. The candidates are repositories on GitHub, found by searches the campaign wrote down before running them; nothing reaches a ledger by another route. Screening follows the campaign's written rules, and the ledger keeps everything — each candidate seen, admitted or excluded, and why. Then the record is sealed.
There are 81 campaigns on the record; 14 of them have had their ledgers read into the survey's two hubs. Ingested means the campaign's admitted ledger rows were extracted for this site — a campaign can be ingested and have admitted no repositories — and the category pages show which campaigns land here. Twelve campaigns predate the current standard; they carry a pre-standard flag, keep their own stated coverage ceilings, and contribute no entries.
Coverage claims are bounded
No campaign claims completeness. Each one states how far it looked and where it stopped, and that wording travels with every entry from that campaign.
The weekly design
The weekly run has not started; its design is below, and the first snapshot is dated 2026-08-10.
Keeping the map current takes three different jobs. Measure: the full cohort is re-pulled weekly from the public GitHub API, by machine, once collection starts. Delta-scout: a short weekly search per task, seeded with the repository ids already held, logging net-new candidates only, screened under the campaign's original rules and appended as dated events. Blind refresh: one rotating task per week across the survey's landed tasks — at the current count, a full blind pass takes more than a year. The blind arm searches its task without the exclusion list — it cannot see what the map already holds — and its finds are merged by numeric repository ID. How much of the standing map it rediscovers is the survey's only internal recall signal. Treating that overlap as a recall estimate rests on assumptions the survey records but does not certify: that the two arms search independently of each other (independence), that every repository is equally findable (homogeneous catchability), and that the field does not change while both look (closure). The recall-calibration campaign's own claim, quoted on its entries, draws the same line: no single overlap certifies recall.
The weekly claim, in full: "cohort re-measured as of <date>; X new candidates screened, Y admitted; map last blind-refreshed <date>" — never "still complete."
Part II
The standard
What an entry page may say, stated in full.
The Persian Cyclamen, from The Temple of Flora, Robert John Thornton (ed.); painted by Philip Reinagle and others, 1799–1807. Amgueddfa Cymru — National Museum Wales (image 132040). Public domain. Credits
What every entry shows
Identity, as metadata facts. The campaign that admitted it, with its coverage claim quoted word for word — never upgraded. Metrics with their retrieval date. Trust marks: snapshot date, snapshot hash, method version, and the sealed record's name. Our own description, labeled as an editorial draft.
What no entry will ever show
Rankings. Repository-authored text — no README excerpts, no self-descriptions, no logos; metadata facts only. Quality language beyond the quoted claim; a build check fails the site if a page tries. Identity lists — watcher and star figures appear as counts, because small denominators identify people. Default order is by category and name, never by popularity.
Some quoted claims contain a campaign's own comparative judgments. The quotes stay unaltered; the site adds none of its own, and nothing here is ordered by them.
What the hash covers
The hash proves the bytes, and only the bytes: what we saw, and when — not whether the search was complete.
The published sha256 covers the raw API capture shared by both hubs of this survey; the capture artifact itself is not yet published. Each hub's served entries file carries its own hash in /api/snapshot.json.
Absence renders
A number we do not have renders as a stated absence — a dash, or a short sentence naming why — never a zero. Repositories that stop resolving stay on the map as dated events — one 404 is an observation, not a death certificate. Weekly movement begins with the second snapshot; until then every movement line says so.
Star histories await a one-time backfill from an independently collected archive of the same platform's public events; the chart's source label is fixed before the chart exists.
Claims carry their own shorthand
Claims sometimes carry campaign-internal shorthand — stop-rule codes, reviewer counts. We quote them anyway, because translation is where upgrading sneaks in; the glossary in Part III covers the recurring terms.
Part III
Provenance and caveats
What the labels do not tell you.
The Maggot-bearing Stapelia, from The Temple of Flora, Robert John Thornton (ed.); painted by Philip Reinagle and others, 1799–1807. Amgueddfa Cymru — National Museum Wales (image 132005). Public domain. Credits
Dated and hashed is not the same as complete. Six caveats travel with every page here, and each names a limit that is still open.
Glossary
The recurring vocabulary, defined once. Campaign-internal shorthand is named here and quoted where it appears; it is not decoded, because decoding a sealed record restates it in this site's voice.
- Campaign
- One task, searched once, on the record: a question and scope frozen first, stated inclusion rules, a screening pass, and a sealed ledger. The campaign is the unit of selection.
- Ledger
- The sealed record a campaign writes: every candidate seen, the disposition it received, and the reason. Ledgers are not public while the licensing review is open.
- Bundle
- One campaign's ledger and its attachments as a single extractable package. The word is used only where a figure is counted per package rather than per campaign.
- Entry
- One repository admitted by at least one campaign ledger and resolved against the dated snapshot. There is one entry per distinct repository on this hub; the numeric repository id is the key where the repository resolved, and the entries that did not resolve are keyed by slug. An entry is whatever a campaign's ledger admitted as evidence: tools, but also specifications, datasets, benchmarks and documentation repositories. The unit is the repository, not the product.
- Cohort
- The distinct repositories the weekly design will re-measure — about 1,246 of them. It is larger than the entry set, because it includes rows no campaign admitted.
- Snapshot
- One dated capture of the public GitHub API, hashed at capture. Every metric on this site carries the date of the snapshot it came from.
- Ingested
- The campaign's admitted ledger rows were extracted for this site. A campaign can be ingested and have admitted no repositories.
- Extracted
- A campaign whose sealed ledger was opened and read into this site's data. Fifteen of the 81 campaigns were selected for extraction, and the rule that picked them is stated on the data card.
- Run type
- The campaign's own designation for how it searched, reproduced as the campaign recorded it. It is a label a campaign applied to itself; this site assesses nothing.
- Claim
- A campaign's coverage statement, quoted verbatim and carrying its own as-of date. Claims are never paraphrased and never upgraded.
- Effort-bounded
- A campaign's term for a search that stopped at a stated effort ceiling rather than at saturation. It bounds the search, not the field.
- Stop-rule codes
- The campaign-internal notations recording why a search stopped — for example K=2 unmet, clocks 0/2, or STOP_RESOURCE_CEILING_OPEN_TAILS. They are named here and quoted where they appear; they are not decoded here.
- AWAITING_INDEPENDENT_QA
- A flag inside some sealed claims: the campaign published with its QA pass pending; the flag stays in the quote until a sealed update clears it.
- scope_id
- The identifier of the frozen scope a campaign searched inside, so two campaigns can be told apart by their boundaries rather than by their dates.
- SCOPE.md
- The file in which a campaign wrote its question and its boundaries down before searching; naming it records that the scope was fixed in advance rather than described afterwards.
- candidate_bytes_acquired
- A campaign-record field recording whether the campaign retrieved the candidate artifacts themselves rather than only their metadata; where it reads false, the campaign screened without acquiring them.
- coverage_target
- The coverage a campaign set out to reach, recorded before it searched. It states an intention, never a result.
- coverage_claim_ceiling
- The strongest coverage statement a campaign allows itself at its stop point; the claim may be read up to it and never past it.
- coverage gate NOT MET
- A campaign's own record that it stopped before its coverage target was reached; what it found still stands as recorded, and no completeness follows from it.
Campaign-internal codenames (reviewer personas, console names) stay unexplained by design; they are part of the sealed wording.
-
Editorial drafts
Descriptions are curator-authored editorial drafts, pending external review, and are labeled as such on every entry.
-
First snapshot
All metrics derive from the first snapshot, 2026-08-10, hash published on every entry. Deltas require a second snapshot; until then, movement renders as a stated absence.
-
Star history pending
The star-history backfill from an independently collected archive of the same platform's public events (reconstructable to February 2011) has not run. When it ships, its source label reads: that archive plus our snapshots.
-
Licensing under review
The underlying database is under a licensing review and is intended to be opened. Until the review lands, pages are quotable and the database is unreleased.
-
Taxonomy versioned
The category map is versioned data (v3.3). Reassignments land as dated events; the mapping is never silently rewritten.
-
Method: supervised beta
The method is a supervised beta, not a ratified standard, and this page says so. Claims are capped by the method's own validation record (sealed with the method, not public).
What is not yet public
The sealed campaign records hold everything you would need to repeat a campaign: the search strings, the inclusion rules, the screening logs, the reviewer notes. None of it is public yet; releasing it is part of the licensing review. Until that lands, you can read how the machinery works on this page and still not be able to run it yourself. That is a real limit on what the survey can be held to, and it is meant to be temporary.
Non-accepted areas
Each area below is recorded as non-accepted: not covered here, and coverage should not be inferred from a neighboring theme.
| Area | State | Detail |
|---|---|---|
| Rankings | Refused | No ordering on this site encodes quality. Category and name only. |
| Repo-authored text | Refused | No README excerpts, project descriptions, or logos appear anywhere. Metadata facts only. |
| Identity lists | Refused | Counts only. No stargazer, watcher, or contributor identities, ever. |
| Peer review | Not performed | Nothing here has been peer-reviewed by us, and presence implies no such thing. |
| Replication | Not performed | We did not run, rebuild, or reproduce any repository's claims. |
| Verification | Not performed | Functional verification was not performed. The survey records what ledgers admitted, not what works. |