mrkeyoor.com_
Wed 07 Oct 14:59 UTC
Self-Hostedevaluationupdated 07 Oct 2026

sparkDash review

sparkDash is a browser dashboard for monitoring NVIDIA DGX Spark machines and other Linux hosts with NVIDIA GPUs. It collects hardware, network, storage, LLM-server, ComfyUI, and optional Hermes Agent status from the local host or over SSH, with controls for benchmarks and power actions.

Verdict

Our sparkDash run built in 7 seconds and passed all 151 Vitest tests, but npm audit found 3 critical and 2 high known vulnerabilities, so the code works while the dependency risk still needs an owner. It is a good fit for a trusted DGX Spark lab that values LLM-aware monitoring and accepts privileged host access. Keep it on loopback or behind authenticated access, and resolve the audit findings before treating it as a shared operations console.

We ran it

Lab card: what happened when we ran sparkDashScreenshot of sparkDash (mia-ai.net)
Install✓ · 6s226 packages · 181 MB
Build✓ · 7s
Tests✓ · 27s151 passed · 0 failed of 151 (vitest)
Known vulns53 critical · 2 high · 0 moderate · 0 low (npm audit)
Repo224 files~45,468 lines of source · 10 MB · 0 CI workflows · Dockerfile

Answers from our run

Does sparkDash build from source?

Dependencies installed in 6 seconds (226 packages), and the build succeeded in 7 seconds. We cloned commit 6e2a394 into a clean Debian container with 3 CPUs and no project-specific setup.

Do sparkDash's tests pass?

Yes: 151 of 151 passed when we ran the project's own test command (vitest). Some failures need services or credentials a bare container does not have.

Does sparkDash have known vulnerabilities in its dependencies?

npm audit flagged 5 known advisories in the dependency tree, including 3 critical at the time of our run.

Who should not use sparkDash?

Operators who cannot run a privileged container with host networking, the host PID namespace, and read-only mounts of /proc, /sys, and the host root.

What are the alternatives to sparkDash?

NVIDIA DCGM Exporter, Netdata, Glances. Our sparkDash run built in 7 seconds and passed all 151 Vitest tests, but npm audit found 3 critical and 2 high known vulnerabilities, so the code works while the dependency risk still needs an owner.

Setup3/5Fast install, but ARM64 and privileged host access narrow deployment
Docs5/5Detailed setup, security, collectors, APIs, and failure notes
Community3/5548 stars, 22 issues and PRs, and an October 6 push
Maturity3/5151 tests pass, but no CI workflows and 5 serious advisories

Who it’s for

DGX Spark owners who want several units visible in one purpose-built dashboard.
Home labs running local LLM servers, ComfyUI, or Hermes Agent on NVIDIA hardware.
Operators comfortable granting a container host-level access for local collectors.
Small GPU fleets that prefer direct SSH polling over installing an agent on every node.

Who it’s NOT for

Operators who cannot run a privileged container with host networking, the host PID namespace, and read-only mounts of /proc, /sys, and the host root.
Public or shared networks without an authenticated proxy or SPARKDASH_TOKEN: the README says an open remote bind lets reachable users change settings and power units off.
Teams with a zero-critical-vulnerability release gate: our npm audit found 3 critical and 2 high known vulnerabilities.
Mixed-architecture deployments expecting the supplied production image to work unchanged: the Compose file targets linux/arm64, and issue 124 asks how to run it on a non-ARM external machine.
Organizations that require repository CI before accepting changes: our scan found 0 CI workflow files.

Setup reality

Our sandbox installed commit 6e2a394 in 6 seconds, adding 226 packages and using 181 MB. The build passed in 7 seconds. Tests passed in 27 seconds, with Vitest reporting 151 passed and 0 failed. Npm audit found 5 known vulnerabilities: 3 critical and 2 high.

The quick Docker route needs an ARM64 host, privileged mode, host networking, the host PID namespace, and mounts into /proc, /sys, /, and NVIDIA libraries. Remote nodes need SSH credentials; LLM and ComfyUI probes need network access to their configured ports.

Loopback is the safe default. Remote access needs an SSH tunnel, Tailscale Serve, an authenticated TLS proxy, or a bearer token. The app can shut down machines, so exposing port 5555 without authentication is a control-plane risk, not only a telemetry leak.

One dashboard understands DGX Spark and local LLM servers

sparkDash watches the things a generic host dashboard tends to miss on an NVIDIA DGX Spark. It shows the shared-memory split, GPU processes, thermal throttling, LLM request rates, model-server queues, and several inference backends. A single view can cover local and remote Sparks, plus ordinary Linux workstations with NVIDIA cards. Remote collection happens over SSH, so there is no sparkDash agent to install on every monitored box.

The checked repository is compact for that scope: 224 files, about 45,468 lines of source, and a 10 MB checkout. React 19 and Vite provide the browser interface, while an Express 5 server runs the collectors, REST endpoints, and WebSocket stream. Version 1.8.9 recognizes llama.cpp, vLLM, SGLang, ds4-server, EXL3, TensorFold, and q27 endpoints. It also has opt-in ComfyUI and Hermes Agent checks.

The specialization is the selling point. Decode and prefill benchmarks can target local or remote OpenAI-compatible servers, and the app understands metrics such as token rates, time to first token, queue state, and KV cache use when a backend publishes them. Results remain dependent on what each server exposes. The README explicitly says TensorFold may show 0 tokens per second when its health response lacks cumulative token totals.

The small Node install hides a powerful host container

Our npm install added 226 packages in 6 seconds and occupied 181 MB. Those are modest numbers beside many dashboard projects. The production Compose file is less modest: it targets linux/arm64, uses host networking, joins the host PID namespace, runs privileged, and mounts host /proc, /sys, and / paths read-only. It also mounts nvidia-smi and driver libraries as a fallback.

Those permissions have a purpose. Local GPU processes and host metrics are otherwise difficult to see from a container, and a loopback-bound LLM server is unreachable through ordinary bridge networking. Still, the dashboard becomes a sensitive operations component. It can inspect machines over SSH, store encrypted SSH passwords, run Hermes updates, cancel ComfyUI work, and invoke a configured shutdown helper. Treat its configuration directory and secret key like administrative credentials.

What happened when we ran it

Our unprivileged Debian sandbox installed commit 6e2a394 in 6 seconds on 3 CPUs with 8 GB of RAM and no secrets. The production frontend built successfully in 7 seconds. Tests finished in 27 seconds, and Vitest reported 151 passed with 0 failed. That is a clean repository-level result for the path we ran. We did not attach a DGX Spark, exercise SSH collection, or measure monitoring accuracy.

The dependency audit was the clear exception. Npm audit found 5 known vulnerabilities: 3 critical and 2 high, with none rated moderate or low. The measurement block does not identify the affected packages or whether a vulnerable path is reachable, so we cannot name a safe workaround from this run. An operator should reproduce the audit, trace each advisory into the runtime or build tree, and document the upgrade decision before exposing the service.

The repository included a Dockerfile and Compose file, but our scan found 0 CI workflow files and no tests directory. The 151 passing tests still count. They run through Node's built-in test runner and Vitest from package scripts, including co-located test paths under server code rather than a root test folder. The missing CI configuration means GitHub does not show us an in-repo workflow that enforces those checks on every proposed change.

Remote access is safe only when you choose the safe mode

The default listener binds to 127.0.0.1:5555, which is sensible. For another computer, the README recommends an SSH tunnel, authenticated reverse proxy, or Tailscale Serve. A direct non-loopback bind can use SPARKDASH_TOKEN. Without that token, remote access is open by default unless SPARKDASH_ALLOW_OPEN_REMOTE=0 forces startup to fail. Reachable users can read telemetry, edit settings, and trigger power actions.

Version 1.8.9 documents the warning plainly, and our lab found 5 high-or-critical audit findings that add another reason to avoid casual exposure. The browser stores its bearer token locally and sends it for protected requests and live telemetry. That is adequate for a trusted small network when transport and access are controlled. It is not a substitute for the identity, audit logs, or role separation expected from a large shared operations platform.

Active fixes do not replace a release process

GitHub showed 548 stars and 22 combined issues and pull requests on October 7, 2026. The repository was pushed on October 6, and the recently updated queue contained feature and fix pull requests. There was no latest GitHub release returned by the API, even though the package and README identify version 1.8.9. Users therefore need to pin a commit or image source instead of relying on a tagged release artifact.

Issue 73 documents excessive SSH session churn in an older remote polling path; it is closed, and the current README describes connection reuse. Issue 124 remains open and asks about running the server from a non-ARM mini PC. Combined with the ARM64 Compose target and 0 CI workflows, that makes the supported shape clear: sparkDash is strongest on the hardware it was built around. On a trusted DGX lab, its specific LLM and fleet controls earn the setup. Elsewhere, DCGM Exporter or Netdata will usually fit the operating model better.

Alternatives

ProjectWhat it isPick it when
NVIDIA DCGM ExporterNVIDIA's GPU metrics exporter for Prometheus and DCGM.pick this instead when you already operate Prometheus and need standard GPU metrics rather than Spark-specific controls.
Netdata gh↗A general system observability agent and dashboard with broad host coverage.pick this instead when mixed infrastructure and general server health matter more than local LLM probes.
Glances gh↗A lightweight cross-platform system monitor with terminal and web views.pick this instead when you need quick host monitoring without Spark power controls or LLM benchmarks.

What people are saying

  1. [github-trending] MiaAI-Lab/sparkDash

Sources

  1. sparkDash repository
  2. sparkDash README at tested commit
  3. sparkDash production Compose file
  4. Remote external-machine issue 124
  5. SSH polling issue 73

More self-hosted reviews

workbuddy2api-panel · jeff · OpenGFW · splash · quivr · bindery · the whole board →