mrkeyoor.com_
Mon 28 Sept 19:26 UTC
AI Toolsevaluationupdated 26 Aug 2026

scientific-agent-skills review

Scientific Agent Skills is a library of instructions, reference material, and scripts that teaches coding agents how to use scientific software and databases. It works with Claude Code, Codex, Cursor, and other Agent Skills hosts, covering research tasks across biology, chemistry, medicine, engineering, and scientific writing.

+697stars / 7d
Verdict

Our build finished in 8 seconds, but the 14-second test run could not collect 105 skills together because their script modules collide, so buyers should validate the exact subset they install. Scientific Agent Skills is useful as a reviewed notebook shelf for agent-assisted research, especially in Claude Code and Codex. It is unsafe as an unquestioned authority: pin versions, inspect scripts, verify live APIs, and keep qualified humans responsible for scientific and clinical conclusions.

We ran it

Lab card: what happened when we ran scientific-agent-skillsScreenshot of scientific-agent-skills (k-dense.ai)
Install✓ · 99s120 packages · 471 MB
Build✓ · 8s
Tests✗ · 14sran, no count parsed
Known vulns0(pip-audit)
Repo2446 files~221,698 lines of source · 277.8 MB · 5 CI workflows · tests dir

Answers from our run

Does scientific-agent-skills build from source?

Dependencies installed in 99 seconds (120 packages), and the build succeeded in 8 seconds. We cloned commit 36d8f13 into a clean Debian container with 3 CPUs and no project-specific setup.

Do scientific-agent-skills's tests pass?

The test command failed in our container, and its output did not report a pass or fail count.

Does scientific-agent-skills have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use scientific-agent-skills?

Anyone expecting the agent to make clinical decisions: the README limits health skills to research, aggregate analysis, and drafts for qualified review.

What are the alternatives to scientific-agent-skills?

Anthropic Skills, K-Dense BYOK, Superpowers. Our build finished in 8 seconds, but the 14-second test run could not collect 105 skills together because their script modules collide, so buyers should validate the exact subset they install.

Setup3/5Easy host install, followed by per-skill packages and credentials
Docs4/5Detailed catalog, safety boundaries, examples, and test commands
Community4/5Recent release and current corrective contributions
Maturity3/5Broad catalog with API drift and awkward repo-wide testing

Who it’s for

Researchers who already use a coding agent and want reviewed starting points for specific scientific tools.
Labs prepared to install only the skills they need and audit every instruction and script first.
Claude Code, Codex, or Cursor users who can manage Python packages, API credentials, and scientific validation.
Developers building repeatable research workflows from public databases and established Python libraries.

Who it’s NOT for

Anyone expecting the agent to make clinical decisions: the README limits health skills to research, aggregate analysis, and drafts for qualified review.
Teams that install agent instructions without inspection: the project's own security warning says skills can execute code, install packages, make network requests, and modify files.
Users who want one credential-free bundle: individual skills may require API keys, approved network domains, or large scientific packages.
Labs unable to pin and revalidate changing APIs: pull request #228 corrects a wrong-strand safety instruction against a live genomics API, and #201 fixes nonexistent Treasury fields.
Operators who expect one repo-wide pytest command: our run stopped during collection because 105 skill script directories share Python module names.

Setup reality

Our sandbox installed 120 packages in 99 seconds and used 471 MB. The build succeeded in 8 seconds. Tests exited with code 4 after 14 seconds because the runner refused to collect 105 skills in one process: their scripts/ folders reuse module names such as _common.py, so each skill must be tested separately.

Installation into a supported host can be one npx skills add or gh skill install command. Repository tooling requires Python 3.13+ and uv; individual skills add their own packages, API keys, services, data downloads, and network permissions.

Do not install the whole catalog by default. The README recommends a topical subset, reading each SKILL.md, scanning scripts, checking contribution history, and pinning a tag or commit.

The catalog gives agents scientific operating instructions

Scientific Agent Skills does not provide a new model or one scientific program. Each skill is a SKILL.md package that tells a compatible agent how to approach a tool or workflow, often with reference files and executable scripts. The catalog spans Python libraries, public databases, lab platforms, literature work, visualization, regulatory evidence preparation, and document production. Claude Code, Codex, Cursor, Gemini CLI, and other hosts can discover the same basic format.

That approach is useful because agents often know a library's vocabulary while missing the safe order of operations, current parameter names, or domain checks. A good skill can put those details near the moment of execution. It cannot make a language model scientifically reliable by itself. The README repeatedly confines clinical material to research, aggregate analysis, or drafts for qualified review, and it rejects patient-specific diagnosis and treatment decisions.

Installing 1 skill is safer than loading the full catalog

The shortest install uses npx skills add K-Dense-AI/scientific-agent-skills. GitHub CLI 2.90.0 or newer can install one named skill and pin a release or commit. The repository also follows the Agent Plugins layout, so plugin-capable clients can discover every immediate child under skills/. Manual copies work when a host scans the chosen user or project directory.

The README itself advises against installing everything. It says community contributions may not have received the same depth of internal review as K-Dense-authored material. A skill can tell an agent to run code, modify files, install packages, or contact external services. That makes each installed folder part documentation, part executable supply chain. Read the instructions and scripts, inspect their history, run a scanner, and pin a version before letting the host use them.

A topical subset also reduces standing context. The README notes that 161 skills add up to a large amount of material, even though the current badge and introduction say 163. Those inconsistent counts are minor as product facts, though they show why the directory itself should be treated as authoritative. Install the 2 or 3 skills for the current project rather than treating the catalog as one indivisible plugin.

What happened when we ran it

Our sandbox installed 120 Python packages in 99 seconds and used 471 MB on disk. The package build succeeded in 8 seconds. pip-audit reported 0 known vulnerabilities. The measured checkout contained 2,446 files and about 221,698 lines of source, with 5 CI workflow files and a tests directory. It had no Dockerfile.

The test command failed during collection with exit code 4 after 14 seconds. Its message was specific: 105 skills could not be collected in one process because their scripts/ directories reuse module names such as _common.py. Python could resolve one skill's helper while testing another. The runner instructed users to run one skill at a time. No test assertions ran in that invocation, so this was a test-layout failure rather than evidence that 105 skill behaviors are broken.

The README's contribution section matches that constraint. It gives a repo-level metadata test, then a separate pytest tests/<skill-name> command for each skill. Contributors can work that way, but a buyer auditing a broad subset must create an isolated loop or rely on CI jobs that preserve module separation. A single green build in 8 seconds says the package assembles; it does not validate scientific results or live API compatibility.

Live APIs can make correct-looking guidance wrong

Two open pull requests show the maintenance problem clearly. Pull request #201 says the U.S. Fiscal Data skill documented Treasury auction fields that do not exist, causing an error or an empty column. The correction replaces the field names and explains that one value described as a percentage is a currency amount. A researcher following the old example could get no result or misread the returned quantity.

Pull request #228 is more serious. It updates the Genomic Intelligence skill after checking the live /v1 contract. The old guidance suggested a wrong-strand splice call could be detected from empty or zero output. Live checks showed the reverse complement could still return high-looking scores, so that heuristic could let an incorrect result pass. The proposed text now requires getting sequence orientation right before the call.

These are healthy corrections, and the repository was pushed on August 24, 2026, with issue and pull-request activity continuing the next day. They also prove that prose can drift even when its format validates and its scripts parse. Any skill that wraps a remote schema needs a fixture against the live service, a dated compatibility note, and an independent check of scientific semantics.

Credentials and scientific responsibility stay with the operator

Repository tooling needs Python 3.13 or newer and uv. Individual skills may broaden Python support, but they can bring large packages, model weights, external databases, and API credentials. A default-deny agent sandbox such as NemoClaw also needs explicit domain approval before package installation or API calls work. The skill file should identify these needs; the operator still has to scope secrets and network access.

For research code, inspect intermediate data rather than accepting the final narrative. Confirm identifiers, units, reference genomes, statistical assumptions, and database dates. For lab automation, simulate before touching equipment. For clinical or regulatory material, keep the generated artifact visibly draft and route it to the qualified person named in the workflow.

Scientific Agent Skills is worth using when it saves the first hour of reading scattered package and API documentation. Its 471 MB base install, per-skill test isolation, and current API corrections rule out blind trust. Choose a small pinned subset, test it against known fixtures, and treat the agent's output as work to verify rather than a scientific result.

Alternatives

ProjectWhat it isPick it when
Anthropic Skills gh↗Anthropic's smaller collection of general document and workflow skills for compatible agents.pick this instead when you need maintained general-purpose skills rather than a broad scientific catalog.
K-Dense BYOKA desktop research workspace that packages these skills with model selection and file tools.pick this instead when you want an integrated application and can bring your own model API keys.
Superpowers gh↗A general coding-agent skill set focused on planning, testing, debugging, and delivery habits.pick this instead when software-engineering discipline matters more than scientific domain coverage.

What people are saying

  1. [github-trending] K-Dense-AI/scientific-agent-skills

Sources

  1. Scientific Agent Skills README
  2. Scientific Agent Skills v2.64.0 release
  3. Genomic Intelligence contract correction
  4. Treasury field correction

More ai tools reviews

rizzo-pii · redamon · SkillOpt · awesome-ai-agent-platforms · guizang-yingzao-skill · vdn-minimax-h3 · the whole board →