mrkeyoor.com_
Mon 28 Sept 07:43 UTC
Dataevaluationupdated 28 Sept 2026

polyledger review

PolyLedger downloads Polymarket market metadata and Polygon trade fills into one local DuckDB database. It gives analysts a resumable history that can be queried with SQL or exported to Parquet without operating a database server.

Verdict

Our PolyLedger run installed 48 packages, passed all 41 tests, and reported 0 known vulnerabilities, making it an easy recommendation for historical Polymarket analysis. Use it when a local SQL file and resumable fills are the goal. A live trader needs order-book capture and reorganization handling that this 64-block-delayed indexer explicitly does not provide.

We ran it

Lab card: what happened when we ran polyledgerScreenshot of polyledger (github.com/nahrek/polyledger)
Install✓ · 32s48 packages · 125 MB
Build✓ · 5s
Tests✓ · 24s41 passed · 0 failed of 41 (pytest)
Known vulns0(pip-audit)
Repo26 files~2,233 lines of source · 0.1 MB · 0 CI workflows · tests dir

Answers from our run

Does polyledger build from source?

Dependencies installed in 32 seconds (48 packages), and the build succeeded in 5 seconds. We cloned commit 8c8d731 into a clean Debian container with 3 CPUs and no project-specific setup.

Do polyledger's tests pass?

Yes: 41 of 41 passed when we ran the project's own test command (pytest). Some failures need services or credentials a bare container does not have.

Does polyledger have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use polyledger?

Live trading systems: PolyLedger stays 64 blocks behind the chain head and does not reconcile already-written reorgs.

What are the alternatives to polyledger?

Polymarket Subgraph, Polymarket V2 Indexer, Squid SDK. Our PolyLedger run installed 48 packages, passed all 41 tests, and reported 0 known vulnerabilities, making it an easy recommendation for historical Polymarket analysis.

Setup4/532-second install; only the HyperSync token adds friction
Docs5/5Schema, math, contracts, commands, and limits are explicit
Community3/5624 stars and a recent push, with no open issue history
Maturity3/541 tests pass, but the young project has no tagged release

Who it’s for

Analysts building repeatable Polymarket trade research in SQL, pandas, Polars, or Spark.
Developers who need raw fill events and a human-readable market join in the same local file.
Researchers who expect long backfills to be interrupted and want transactional checkpoints.
Python users willing to obtain an Envio HyperSync token for Polygon event history.

Who it’s NOT for

Live trading systems: PolyLedger stays 64 blocks behind the chain head and does not reconcile already-written reorgs.
Order-book research: quotes, open orders, and cancellations do not appear on chain and are outside this index.
Multi-venue prediction-market analysis: the README says Kalshi needs a different collector, which has not been written.
Users who cannot obtain a HyperSync API token: it is required even though Polymarket's CLOB and Gamma endpoints are public.
Teams expecting a released service with deployment files: the repository has no tagged release, CI workflow, or Dockerfile.

Setup reality

Our sandbox install succeeded in 32 seconds, adding 48 packages and using 125 MB. The build passed in 5 seconds. Pytest passed all 41 tests in 24 seconds with 0 failures, and pip-audit found 0 known vulnerabilities.

Python 3.11 or newer and an Envio HyperSync API token are required. CLOB and Gamma calls need no credentials. A .env file is enough for the token and database path, but the README says a first historical backfill on the free tier can take from hours to days.

The output is one DuckDB file, with optional Parquet exports. The indexer deliberately stays 64 blocks behind the chain head, stores filled trades only, and does not roll back data already written after a reorganization. The repository has tests but no CI workflow or Dockerfile.

PolyLedger puts market names and chain fills in one DuckDB file

Polymarket analysis usually starts with a join problem. The CLOB API knows the question, outcome labels, and token identifiers, while Polygon records the actual OrderFilled events. PolyLedger gathers both sides, decodes the current and older exchange contracts, and writes a ready-to-query trade view into DuckDB. You get human-readable markets beside maker, taker, price, shares, dollar size, fee, block time, and the raw event fields used to derive them.

The codebase is unusually small for that job: 26 files, about 2,233 lines of source, and a 0.1 MB checkout at commit 8c8d731. Its CLI has separate commands for markets, chain events, joined trades, statistics, SQL, and Parquet export. Python 3.11 or newer is required. HyperSync supplies Polygon history, while Polymarket's public CLOB and Gamma APIs provide market metadata and fill gaps.

Rows and checkpoints commit in the same transaction

Resumption is the central design choice. PolyLedger writes each batch and its block cursor together, so a crash cannot leave data committed with an old checkpoint or advance the cursor past missing rows. A unique key on transaction hash and log index makes repeated ranges harmless. V1 and V2 contracts keep separate checkpoints, which lets an analyst add older history without corrupting the current exchange stream.

Network behavior is designed for a long backfill. The client validates API responses with Pydantic, follows Retry-After, adds exponential backoff with jitter, and shares an 8-request-per-second default limiter across REST sources. Missing token IDs remain visible as unmatched fills and can be backfilled from Gamma. The README says the first free-tier sync may take hours or days, so interruption safety is doing real work rather than decorating the feature list.

What happened when we ran it

Our sandbox installed PolyLedger in 32 seconds, adding 48 packages and consuming 125 MB. The build completed in 5 seconds. Pytest then passed all 41 tests in 24 seconds with 0 failures. Pip-audit found 0 known vulnerabilities. For a Python indexer that touches external APIs, ABI decoding, derived arithmetic, and local storage, that is a reassuring first run.

The tests do not need a network connection. According to the README, fixtures cover V1 and V2 decoding, duplicate inserts, transaction rollback, price and side math, pagination, a stuck-cursor guard, batching, rate limits, and retry timing. Our repository scan found a tests directory but 0 CI workflow files and no Dockerfile. The suite is present and healthy at the measured commit, though GitHub does not show it being enforced on every push.

Each trade keeps both sides of the fill

Price and share quantity are derived from the maker and taker amounts rather than stored directly on chain. PolyLedger records maker_side and taker_side, avoiding the ambiguous single side field common in exported trade datasets. It also retains raw fills in their own table. If you disagree with the interpretation, you can rebuild the join and arithmetic from source columns instead of accepting a transformation you cannot inspect.

That transparency matters across 2 contract generations. V2 emits the side explicitly. V1 requires inference from which asset leg is collateral, and token-for-token fills without a collateral leg are skipped because no useful price can be calculated. Example SQL covers daily volume, candles, VWAP, positions, and top traders. Materializing the joined view is optional for repeated queries, while Parquet export moves every table into other analytical systems.

A 64-block delay keeps this out of the trading loop

PolyLedger protects historical collection by stopping 64 blocks behind Polygon's head. It avoids writing fills from blocks most exposed to reorganization, but it does not reconcile rows already stored if a deeper reorg occurs. That contract is suitable for research and explicitly insufficient for live trading. A system placing orders should maintain a canonical-head process and repair orphaned events rather than treating the local file as immediate truth.

The data scope has another firm edge. Only filled trades reach this database. Quotes, open orders, and cancellations live in the CLOB order book and require a real-time WebSocket collector. Separate FeeCharged events are also omitted, with fees taken from the fill event field. Kalshi is absent because its settlement path needs a different collector. These omissions are documented, which makes it easier to decide without mistaking a 125 MB install for a complete market-data platform.

September activity is promising, with no release history yet

GitHub recorded a last push on September 26, 2026, 624 stars, and 0 combined open issues and pull requests. The repository was created on September 2 and has no tagged release. That is current activity, not evidence of long operating history. Pin a commit, keep the DuckDB file backed up, and run the 41-test suite before adopting schema or decoder updates.

The official Polymarket subgraph is a better fit for GraphQL and its wider set of activity, PnL, liquidity, and order-book entities. Envio's Polymarket V2 Indexer covers fees, pUSD flows, collateral operations, rewards, and aggregates, at the cost of a larger service stack. Squid SDK fits teams designing a custom multi-chain ETL system. PolyLedger wins when the desired artifact is simply a local SQL file that can resume.

Historical Polymarket research is the right job

PolyLedger has a narrow promise and meets it cleanly in our run. Forty-one passing tests support the decoder and transaction claims, while 0 audit findings lower the cost of a trial. The required HyperSync token is the only external gate before metadata and chain collection begin.

Keep it out of execution paths that need the current book. For notebooks, reproducible reports, and address-level fill studies, the atomic cursor and visible unmatched rows are exactly the controls an analyst needs. The decision turns on one question: can your answer tolerate being 64 blocks behind? If yes, this is the first project in the batch I would install without much hesitation.

Alternatives

ProjectWhat it isPick it when
Polymarket SubgraphPolymarket's Graph-based indexers for activity, order-book, PnL, liquidity, and other datasets.pick this instead when GraphQL and broader indexed entities matter more than a single local DuckDB file.
Polymarket V2 IndexerAn Envio HyperIndex project covering V2 fills, fees, transfers, collateral flows, rewards, and aggregates.pick this instead when you need more V2 event types and are willing to operate Docker or Podman plus a GraphQL service.
Squid SDKA TypeScript toolkit for building custom blockchain ETL pipelines across several chain families.pick this instead when Polymarket is one source in a larger custom multi-chain indexer.

What people are saying

  1. [velocity-scout] nahrek/polyledger

Sources

  1. PolyLedger README
  2. PolyLedger project metadata
  3. PolyLedger example queries
  4. PolyLedger source repository

More data reviews

timeseries-atlas · opendataloader-pdf · data-formulator · toasty · gfwlist · simdjson · the whole board →