Eight extractors turn local agent history into JSONL
AI Data Extraction searches common application folders for Claude Code, Codex, Cursor, Trae, Windsurf, Continue, Gemini CLI, and OpenCode records. Each extractor understands some mix of JSON, JSONL, SQLite, or application-specific data. It writes timestamped JSONL files with a normalized conversation shape. Depending on the source, records can include messages, code selections, suggested edits, tool calls, results, project paths, model names, and token use.
That breadth solves a real migration problem. Coding assistants keep useful history in different folders and schemas, often without a shared export format. A normalized file is easier to search or feed into an authorized data pipeline. It is also more dangerous: one output directory can concentrate prompts, proprietary source, shell output, tokens, personal paths, and model reasoning that were previously scattered across 8 applications.
The 19-second setup is the easy part
Our sandbox installed 35 Python packages in 16 seconds and used 37 MB. The build passed in 1 second. Pytest then passed 6 of 6 tests in 2 seconds, and pip-audit found 0 known vulnerabilities. At commit cc36017, the repository contained 13 files, about 3,359 lines of source, and occupied just 0.1 MB.
Those numbers make source review feasible. They do not validate every supported assistant version. Cursor alone has several storage layouts in the README, while OpenCode spans CLI JSON, desktop data, and a newer SQLite format discussed in an open pull request. Our scan found a tests directory but 0 CI workflow files and no Dockerfile. The tests passed locally in our harness; GitHub was not visibly running them on every change.
What happened when we ran it
Our run installed the project in 16 seconds, built it in 1 second, and completed tests in 2 seconds on 3 CPUs with 8 GB of RAM. All 6 tests passed, with 0 failures, and the installed dependency audit reported 0 known vulnerabilities. No setup failure appeared in the supplied results.
We did not point the scripts at a real home directory or export private assistant sessions. That means the lab result covers packaging and the available tests, not extraction completeness, secret detection, or compatibility with each live application. We also did not call the optional skill-generation path or send a corpus to any model endpoint. Those actions require user data and a privacy decision that a generic sandbox should not make.
Raw export happens before the proposed privacy filters
The README openly warns that extracted files may contain proprietary code, API keys, secrets, and personal file paths. Its current advice is to run detect-secrets, review the files, keep them on encrypted storage, and avoid committing them publicly. Those are useful warnings, but they are operator instructions rather than an enforced safe default. The main extraction path still materializes the raw corpus first.
Two open pull requests make the gap easier to see. Pull request 21 proposes nested-string redaction with a local privacy model. Pull request 18 proposes a stricter Claude Code export that drops tool calls, diffs, and unrecognized content by default, then records a manifest and audit warnings. Neither was merged when we checked. Do not describe those branches as shipped protection, and do not build a policy around code that is still under review.
Corpus-to-skill conversion creates a second disclosure boundary
The included corpus_to_skills.py samples conversations and asks an OpenAI-compatible chat-completions model to produce Agent Skills. By default it targets a local endpoint at 127.0.0.1:8000; setting OPENAI_BASE_URL can send the sample elsewhere. The README tells users to filter the corpus before configuring a remote API because generated output can repeat sensitive source material.
That warning should be treated as a release gate. Review the exact sampled input, endpoint, provider retention policy, generated SKILL.md, and destination repository. A secret scanner will not reliably classify confidential algorithms, customer names, internal URLs, or business rules. The 6 passing tests cannot establish that a model will avoid reproducing such content. Human approval is still required before installing or sharing a generated skill.
Storage formats will keep moving under the extractors
The repository was pushed on August 19, 2026. GitHub listed 1,256 stars and 14 combined issues and pull requests. Issue 22 and its pull request sought Grok Build support on August 23, while pull request 8 remained open for OpenCode v1.2.0 SQLite storage. That activity shows demand and adaptation, but it also confirms that compatibility follows upstream changes.
No GitHub release exists, and the repository metadata reported no license. The first point makes commit pinning more important. The second is a legal adoption problem, not missing polish: without permission terms, a company should not assume it can redistribute or incorporate the code. For a personal, local, read-only export, this toolkit is impressively direct. For organizational training data, use it only inside a documented consent, redaction, validation, and retention process.

