Posted by Tsua00021
Jul 22, 2026/15:20 UTC
The recently launched repository on GitHub titled bitcoin-wizard-bench features a variety of resources including 12 seed items, an item schema, a validator, and Continuous Integration (CI) tools. The repository incorporates specific questions from @alexwaltz which are designed to evaluate the authenticity and proficiency of systems in their understanding of Bitcoin, crediting Alex Waltz for the conceptual framework. Two items have been meticulously curated to serve as benchmarks: the first involves exploring discrepancies between block explorers, focusing on block 78's coinbase transaction with a Pay-to-Public-Key (P2PK) output, and contrasts raw-public key renderings against a derived address that does not appear on-chain. Additionally, it includes a label heuristic for miners. The second curated item addresses the SIGHASH_SINGLE quirk, strictly referring to the verification methods detailed in the legacy SignatureHash() function from Bitcoin Core and the footnote in the BIP143 Specification.
The remaining draft items in the repository are intentionally incomplete, encouraging contributions through pull requests (PRs) that complete these entries with well-sourced information. This approach aims to engage the community in developing a robust set of challenges that accurately assess systems' capabilities in cryptographic and transactional nuances of Bitcoin.
Further elaboration on the project reveals an overlap with another initiative led by @brenorb, focusing on dataset tracking rather than benchmarking, ensuring no duplication of efforts. An interesting aspect of the project is the incorporation of a hierarchical structure of constitutional and precedent sources which assigns normative authority to various sources, enhancing the metadata for scoring disagreements within the contested tier. This structure differentiates the weight given to different types of sources, such as a Bitcoin Improvement Proposal (BIP) versus a forum post from 2010, in attributing positions within the taxonomy.
Plans for future development include deploying an evaluation runner for tier 1 and subsequently publishing results from initial runs of frontier models, both with and without supplementary tools. This phased approach underscores a systematic enhancement and validation of the benchmarking framework, aiming to establish a comprehensive and reliable metric for evaluating Bitcoin-related systems.
Thread Summary (21 replies)
Jun 2 - Jul 22, 2026
22 messages
TLDR
We’ll email you summaries of the latest discussions from high signal bitcoin sources, like bitcoin-dev, lightning-dev, and Delving Bitcoin.
We'd love to hear your feedback on this project.
Give Feedback