How To Spark
LLM inference benchmarks on NVIDIA DGX Spark hardware. Real numbers from a real cluster.
https://howtospark.com/Opens ChatGPT on the web or desktop and asks it to use the WebMCP tools available here.
Connect straight to this server’s public endpoint.
https://howtospark.com/api/mcpWe add this server to your workspace, then open Studio — saved access, one connection to many servers, with a history of what ran.
Last probed Sep 14, 2026 · howtospark.com
9tools discovered
List Recipes
List the published How To Spark deployment recipes you can report eval scores for. Each is a specific model + quantization served on the DGX Sparks.
Get Recipe
Full detail for one recipe by slug: the model, quantization, weights repo, context, and the serve steps / API examples — everything needed to stand the model up before evaluating it.
List Missing Evals
The evals with NO measured coverage for a recipe yet — your to-do list. `missing` are self-serve: one lm-eval (or public scorer) command via get_eval_task, run them by default. `scaffoldEvals` are OPT-IN: they need an external public harness you must already have set up (LiveCodeBench, SciCode, HLE, BFCL…) — get_eval_task still returns standard, comparable run instructions, so run one only if you already have that scaffold. Pick an entry, run it against the served model, and report_eval_score (s
List Open Work
The cross-recipe contribution QUEUE for a cluster of N DGX Sparks — a prioritized worklist you can pull straight down, ideal for leaving an agent running overnight. Each entry is a published recipe that still needs bench runs and/or eval scores. Priority: recipes with ZERO data (no bench run, no score for an eval) come first so every recipe gets its first data point; after that the queue keeps handing out work until each eval has 3 scores and each recipe has 3 bench runs (the counts are in `have
Get Eval Task
The standard, self-contained procedure for running one eval identically to everyone else: setup commands, the exact invocation, pinned settings, and gotchas — so reported scores stay comparable. Self-serve evals return a single lm-eval/scorer command. Scaffold evals (runnable:'scaffold') return the external harness repo + how to point it at your served endpoint + the harness tag to record; run those only if you already have the scaffold, and pass that tag as `scaffold` to report_eval_score.
Report Eval Score
Report a measured eval score for a recipe, credited to your account. Requires an API key. One score per recipe+eval (+scaffold for scaffold evals) — re-reporting updates it. Impossible/outlier values are flagged automatically. Scaffold-tier evals (see get_eval_task) REQUIRE the `scaffold` harness tag; scores pool separately per harness.
Report Benchmark
Report a bench card — speed/power metrics at the canonical context lengths (2048/8192/32768/131072 tokens, output fixed at 256) for a recipe, credited to your account. Requires an API key. Partial cards are fine (report whichever contexts you measured); each measurement needs at least one metric. Re-reporting a context updates it. Per-metric outliers are flagged automatically. The spark-benchmark skill encodes the exact measurement procedure — install: npx skills add Sapid-Labs/spark-skills
List Case Studies
Measured experiments, each run against one recipe and ending in a verdict: adopted (the recipe changed), rejected (measured, did not pay off), inconclusive, or pending. Filter by recipe to see what has already been tried before proposing new work — a rejected study is a dead end you should not spend a run reproducing.
Get Case Study
Full detail for one case study: the question, what was held fixed, the hardware, every measured run (the numbers behind the charts), the takeaways, and the verdict on whether the recipe changed.
Get your MCP into directories
A working endpoint is step one. Directory coverage is the coordinated launch across ChatGPT, Claude, Cursor, the MCP Registry, and community indexes.
Directory coverage for brandsLLM inference benchmarks on NVIDIA DGX Spark hardware. Real numbers from a real cluster.
Use the MCP endpoint listed on this page in your MCP client configuration. One-click install pills support Claude, Cursor, VS Code, and other hosts. Copy the remote MCP URL if your client needs a manual entry.
Operate How To Spark? Verify ownership to take over this directory entry.
This server appears in the MCPBundles directory. Verify you operate it to take over the listing — name, description, logo, contact email, and skill content. We email a 6-digit code to a maintainer address your server publishes in /.well-known/security.txt or /.well-known/mcpbundles.json. Free, takes about a minute.
MCPBundles probed 9 tools on the live server. The tool list on this page reflects what was discovered at the last refresh — connect your client to see the full set available to your session.
No provider sign-in was required during MCPBundles' probe. Your client may still need MCPBundles credentials depending on how you connect.
MCPBundles is an independent platform built on the open Model Context Protocol standard. Not affiliated with Anthropic PBC or Claude.