Your tracker holds the runs and your registry holds the versions. Neither holds the part you need at the moment precision drops: which corpus this was fine-tuned on, what the relabelling pass touched, why the eval set said it was fine, and what the team concluded last time. Filamental keeps that lineage as a structure, in files your own agents can read.
A model, a prompt revision, an eval, a corpus, a deployment, every week. The tracker has all of the runs and the registry has all of the versions, so nothing is lost in the sense of being deleted. But the reasoning that connected them was never an artefact: why that experiment was worth doing, what it settled, which idea it killed, and which known weakness everybody agreed to accept for now.
Six months later that’s all anybody actually wants. And the honest answer is usually a person, a Slack thread nobody can find, and a notebook whose kernel state is gone. The system is fully reproducible and completely unexplainable.
A registry replaces v3 with v4 and archives the old row, because that is what a registry is for. Here both are nodes and both stay, along with the fine-tuning run between them, the corpus each was trained on, the eval set that scored them and the deployment that serves one of them. The lines are typed and directional and you name the types yourself, so trained on, superseded by, evaluated by and produced are different relationships rather than one grey line meaning "related".
That turns two expensive questions into navigation. What did this come from walks backwards through the runs and the data. What breaks if this changes walks forwards through everything downstream of a corpus you’re about to relabel. Both directions are one click, and neither depends on whoever ran it still being in the company.
The notes carry the part that has nowhere else to live. Not the metrics, which your tracker already has and does better, but the sentence explaining what an experiment settled, and the honest one recording what a benchmark doesn’t cover. An eval set that passes is evidence about the eval set.
A nine-person team selling support-ticket classification as an API. The production model was fine-tuned from the previous version on the first quarter's ticket corpus, passed its evaluation at release, and has been running six points of precision below that benchmark ever since. The weekly regression suite caught it on a live-traffic sample. The golden set never did, because none of its two hundred and forty hand-reviewed examples happened to be the billing edge cases the new model struggles with.
The leading theory is that a batch of billing tickets in that corpus was silently mislabelled during an automated relabelling pass run in May, and there’s an experiment open right now testing exactly that. If it holds, the fix is a partial retrain rather than a rollback, which is a week of difference.
All of which hangs on one question: which tickets did the relabelling pass actually touch. Nobody has a complete record. It exists in the memory of the engineer who ran it, and the whole investigation is currently blocked on that one person's recall of a script they executed three months ago. The node for that pass is where the answer should have been sitting.
There’s no SDK to import, no callback to register, no decorator on your training loop and no service to run. It doesn’t watch a directory for checkpoints or scrape your runs. It is a desktop application that reads a folder of files on the machine it’s installed on, which is also why the whole thing works on a plane.
No account, no sign-in and no telemetry, and that last one is a decision rather than a setting: nothing is collected, so nothing can be requested or leaked. Given what tends to be written in the notes next to a model that’s underperforming, that matters more here than on most pages.
It won’t track experiments, store artefacts, serve a model or tell you a run has finished. Keep the tools that do. This is for the layer above them, which is the one that currently lives in people.
Every other kind of work on this site benefits from structure because people navigate it. You have agents open all day. Filamental ships a local MCP server, so Claude Desktop, Claude Code, Cursor or anything else speaking the protocol can search the space, read a node, follow its relationships and write new ones back, running on your machine against your folder with nothing in between. Ask what the last three experiments on this corpus concluded and it answers from the structure rather than from a context window stuffed with documents.
The same structure is what makes it worth writing at all. A corpus node with two sentences about what its labels actually mean is worth more to an agent than the corpus itself, because that’s the part not recoverable from the data. And when somebody outside needs it, a customer asking how a model is evaluated or a diligence request arriving mid-raise, the space can be sent as a link or a single file that opens in a browser, and they install nothing.
You don’t begin with a blank screen. A Template is a starting vocabulary, the kinds of thing that exist in a job and the ways they relate, so the categories are already there and already coloured when you make a space.
A starting point, not a schema you’re stuck inside. Rename a category, add one, delete the ones you never use. Nothing stops you using more than one in a space. Teams shipping a product rather than a model tend to reach for System Component Map alongside these.
Everything above except the sending is on the free plan, permanently, with no account and no card. There’s no per-user tier, no workspace minimum and no tracked-run allowance, which is worth saying out loud on a page for people whose stack is billed by all three.
Spaces are unlimited and the bridges joining them are free, so lineage becomes a few linked spaces rather than one unreadable one: data and labelling in the first, models and experiments in the second, serving and monitoring in the third, each one click from the others. Each holds twenty nodes, which is roughly where a diagram stops being legible anyway.
The paid tier is $120 a year, or $12 a month, and buys exactly one thing: handing a space to somebody who doesn’t have Filamental. For most teams that is a customer question about evaluation, a diligence pack, or a review by people who won’t install a desktop application.
No, and you should keep them. Those tools capture runs: hyperparameters, curves, artefacts, metrics, automatically and at a volume no human would maintain by hand. What they don’t hold is the reasoning between runs, which is why an experiment was worth doing, what it settled, and what the team concluded and moved on from. Filamental holds that layer and links out to the run rather than duplicating it.
Yes. Filamental ships a local MCP server, so Claude Desktop, Claude Code, Cursor or anything else speaking MCP can search the space, read a node, follow relationships and write new ones back. It runs on your machine against your folder, with no cloud service in between. There is also a skill file you can hand any assistant that doesn’t speak MCP, and the exact text we give your AI is published on the site.
In a folder you choose, as one Markdown file per node with YAML frontmatter. Nothing is proprietary and nothing leaves the machine: there’s no account, no sign-in and no telemetry of any kind. That also means a space is a Git repository like any other, so the record of what the team decided is diffable and reviewable alongside the code it describes.
A diagram is a picture of one moment, and the reason pipeline diagrams go stale is that the pipeline changes weekly while the picture doesn’t. Here every version stays present as its own node, superseded rather than overwritten, so v3 and v4 and the run between them all exist at once. The question a diagram can’t answer is what changed and why, and that’s the question this is shaped around.
Yes. Publisher turns a space into a document you send as a link or a single HTML file, and it opens in an ordinary browser with the structure navigable inside it. Useful for a customer asking how a model is evaluated, a due-diligence request, or an internal review by people who won’t install anything. The recipient installs nothing and pays nothing. One published link at a time is free, and the paid tier, at $120 a year or $12 a month, adds more links, updates to a link already sent, the HTML download and presenting.
The Personal plan is free permanently, with no account and no card, and it isn’t seat-priced, so it doesn’t get more expensive as the team grows. Spaces are unlimited and the bridge nodes joining them are free, so lineage becomes a few linked spaces rather than one enormous one. Each holds twenty nodes. The paid tier is $120 a year, or $12 a month, and buys sending a space to somebody who doesn’t have Filamental.
One investigation is enough to find out whether this suits how you work. Free, no account, no card, and nothing to wire into your training loop.