mrkeyoor.com_
Wed 30 Sept 20:34 UTC
Dataevaluationupdated 26 Aug 2026

ai-data-extraction review

AI Data Extraction is a set of local Python scripts that reads stored conversations from coding assistants and converts them into a common JSONL shape. It targets Claude Code, Codex, Cursor, Trae, Windsurf, Continue, Gemini CLI, and OpenCode, including prompts, replies, file context, edits, tool calls, and metadata when the source format contains them.

+4stars / 7d
Verdict

Our sandbox installed 35 packages and passed all 6 tests in 19 seconds combined, so the code is cheap to inspect and trial. Use it only on data you own or are authorized to process, behind a separate redaction and review step, because the reviewed branch writes sensitive coding history without the protections proposed in two open pull requests. The missing repository license is enough reason for a company to pause before reuse or redistribution.

We ran it

Lab card: what happened when we ran ai-data-extractionScreenshot of ai-data-extraction (github.com/0xSero/ai-data-extraction)
Install✓ · 16s35 packages · 37 MB
Build✓ · 1s
Tests✓ · 2s6 passed · 0 failed of 6 (pytest)
Known vulns0(pip-audit)
Repo13 files~3,359 lines of source · 0.1 MB · 0 CI workflows · tests dir

Answers from our run

Does ai-data-extraction build from source?

Dependencies installed in 16 seconds (35 packages), and the build succeeded in 1 seconds. We cloned commit cc36017 into a clean Debian container with 3 CPUs and no project-specific setup.

Do ai-data-extraction's tests pass?

Yes: 6 of 6 passed when we ran the project's own test command (pytest). Some failures need services or credentials a bare container does not have.

Does ai-data-extraction have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use ai-data-extraction?

Organizations that need a clear reuse license: GitHub reported no repository license when we fetched the project.

What are the alternatives to ai-data-extraction?

Official account exports, Langfuse, OpenHands. Our sandbox installed 35 packages and passed all 6 tests in 19 seconds combined, so the code is cheap to inspect and trial.

Setup5/516-second install and all 6 tests passed
Docs4/5Storage paths, formats, fields, caveats, and privacy risks are explicit
Community3/51,256 stars with active August 2026 issues and pull requests
Maturity2/5Small tested codebase, but no license, releases, CI, or merged redaction

Who it’s for

Individuals exporting their own local coding-assistant history for search, analysis, or an authorized training corpus.
ML teams that can inspect each extractor, validate the normalized output, and run a separate privacy review.
Researchers comparing how several coding tools store sessions across JSON, JSONL, and SQLite.

Who it’s NOT for

Organizations that need a clear reuse license: GitHub reported no repository license when we fetched the project.
Anyone who cannot prove rights and consent for every exported conversation: the README warns that output may contain proprietary code, secrets, and personal paths.
Teams expecting built-in redaction before files are written: privacy-filter pull requests 18 and 21 were still open, so those controls were not in the reviewed main branch.
Users who assume new assistant storage formats work automatically: open pull request 8 adds OpenCode v1.2.0 SQLite support, while the README describes legacy JSON and Tauri data.

Setup reality

Our install succeeded in 16 seconds, adding 35 Python packages and 37 MB. The build passed in 1 second, and pytest passed all 6 tests in 2 seconds. Pip-audit found 0 known vulnerabilities. The checkout was 0.1 MB with 13 files and about 3,359 source lines.

The extraction scripts themselves use local assistant files and may need permission to read application data directories. corpus_to_skills.py can call a local OpenAI-compatible endpoint or a configured remote one; remote use needs its model endpoint and credentials.

There was a tests directory, but our scan found 0 CI workflows and no Dockerfile. Output is raw JSONL in extracted_data/; secret scanning, path sanitization, consent checks, access control, and encrypted storage remain the operator's job.

Eight extractors turn local agent history into JSONL

AI Data Extraction searches common application folders for Claude Code, Codex, Cursor, Trae, Windsurf, Continue, Gemini CLI, and OpenCode records. Each extractor understands some mix of JSON, JSONL, SQLite, or application-specific data. It writes timestamped JSONL files with a normalized conversation shape. Depending on the source, records can include messages, code selections, suggested edits, tool calls, results, project paths, model names, and token use.

That breadth solves a real migration problem. Coding assistants keep useful history in different folders and schemas, often without a shared export format. A normalized file is easier to search or feed into an authorized data pipeline. It is also more dangerous: one output directory can concentrate prompts, proprietary source, shell output, tokens, personal paths, and model reasoning that were previously scattered across 8 applications.

The 19-second setup is the easy part

Our sandbox installed 35 Python packages in 16 seconds and used 37 MB. The build passed in 1 second. Pytest then passed 6 of 6 tests in 2 seconds, and pip-audit found 0 known vulnerabilities. At commit cc36017, the repository contained 13 files, about 3,359 lines of source, and occupied just 0.1 MB.

Those numbers make source review feasible. They do not validate every supported assistant version. Cursor alone has several storage layouts in the README, while OpenCode spans CLI JSON, desktop data, and a newer SQLite format discussed in an open pull request. Our scan found a tests directory but 0 CI workflow files and no Dockerfile. The tests passed locally in our harness; GitHub was not visibly running them on every change.

What happened when we ran it

Our run installed the project in 16 seconds, built it in 1 second, and completed tests in 2 seconds on 3 CPUs with 8 GB of RAM. All 6 tests passed, with 0 failures, and the installed dependency audit reported 0 known vulnerabilities. No setup failure appeared in the supplied results.

We did not point the scripts at a real home directory or export private assistant sessions. That means the lab result covers packaging and the available tests, not extraction completeness, secret detection, or compatibility with each live application. We also did not call the optional skill-generation path or send a corpus to any model endpoint. Those actions require user data and a privacy decision that a generic sandbox should not make.

Raw export happens before the proposed privacy filters

The README openly warns that extracted files may contain proprietary code, API keys, secrets, and personal file paths. Its current advice is to run detect-secrets, review the files, keep them on encrypted storage, and avoid committing them publicly. Those are useful warnings, but they are operator instructions rather than an enforced safe default. The main extraction path still materializes the raw corpus first.

Two open pull requests make the gap easier to see. Pull request 21 proposes nested-string redaction with a local privacy model. Pull request 18 proposes a stricter Claude Code export that drops tool calls, diffs, and unrecognized content by default, then records a manifest and audit warnings. Neither was merged when we checked. Do not describe those branches as shipped protection, and do not build a policy around code that is still under review.

Corpus-to-skill conversion creates a second disclosure boundary

The included corpus_to_skills.py samples conversations and asks an OpenAI-compatible chat-completions model to produce Agent Skills. By default it targets a local endpoint at 127.0.0.1:8000; setting OPENAI_BASE_URL can send the sample elsewhere. The README tells users to filter the corpus before configuring a remote API because generated output can repeat sensitive source material.

That warning should be treated as a release gate. Review the exact sampled input, endpoint, provider retention policy, generated SKILL.md, and destination repository. A secret scanner will not reliably classify confidential algorithms, customer names, internal URLs, or business rules. The 6 passing tests cannot establish that a model will avoid reproducing such content. Human approval is still required before installing or sharing a generated skill.

Storage formats will keep moving under the extractors

The repository was pushed on August 19, 2026. GitHub listed 1,256 stars and 14 combined issues and pull requests. Issue 22 and its pull request sought Grok Build support on August 23, while pull request 8 remained open for OpenCode v1.2.0 SQLite storage. That activity shows demand and adaptation, but it also confirms that compatibility follows upstream changes.

No GitHub release exists, and the repository metadata reported no license. The first point makes commit pinning more important. The second is a legal adoption problem, not missing polish: without permission terms, a company should not assume it can redistribute or incorporate the code. For a personal, local, read-only export, this toolkit is impressively direct. For organizational training data, use it only inside a documented consent, redaction, validation, and retention process.

Alternatives

ProjectWhat it isPick it when
Official account exportsVendor-provided exports preserve the data each service officially exposes under its current account controls.pick this instead when provenance, support, and policy compliance matter more than cross-tool normalization.
Langfuse gh↗A self-hostable tracing system that records model calls and application sessions as they happen.pick this instead when you can instrument future workloads and want governed traces rather than scraping old local stores.
OpenHands gh↗An open coding-agent platform whose own event history can be collected within a controlled deployment.pick this instead when you can standardize on one agent environment and capture authorized trajectories at creation time.

What people are saying

  1. [github-trending] 0xSero/ai-data-extraction

Sources

  1. AI Data Extraction README
  2. OpenAI Privacy Filter redaction pull request
  3. Enterprise-safe Claude Code export pull request
  4. OpenCode SQLite support pull request
  5. Grok Build support request

More data reviews

TradeGenuis-box · awesome-submitlist · ccf-deadlines · instagram-private-graph · OpenBB · polyledger · the whole board →