AI Project Template
A board built around an honest fact about AI work (most experiments don't ship), with a real column for the ones that get shelved instead of quietly disappearing.
Shelved isn't a graveyard. It's a wiki entry explaining what was tried and why it stopped, so nobody re-runs it blind next quarter.
Why most AI project boards quietly become a lie
Most AI work doesn't ship. That's not a failure rate to hide. It's close to the normal shape of the work, and a board that pretends otherwise stops being useful fast. A generic task board assumes most things started will eventually finish and go live; AI experimentation doesn't work that way, and forcing it into that shape produces a board that either lies about progress or gets abandoned. Three habits cause that.
Every idea is tracked as if it's heading to production. A hypothesis and a shipped feature get the same card type, so the board can't tell you, at a glance, whether something is a two-day experiment or a committed roadmap item, and everyone plans as if it's the latter.
Experiments that don't pan out just vanish. Nobody wants to write "we tried this and it didn't work" on a board that looks like a delivery tracker, so the card gets archived, deleted, or simply stops being updated. Three months later, someone tries the same thing again, because there was no record it had already been tried.
"Evaluated" gets skipped entirely. An experiment runs, the numbers come back mediocre, and it either ships anyway because momentum carries it, or it dies without anyone writing down why the results weren't good enough.
The cost isn't abstract. A team that treats every experiment as a future feature ends up either shipping things that shouldn't have shipped, because stopping feels like admitting failure, or quietly burying the ones that don't work, which guarantees somebody tries the same idea again in six months with no memory that it was already tried and found wanting.
This template treats "shelved" as a legitimate, tracked outcome, as normal as shipped, with a place to log what was learned. Nothing here integrates with a model registry or an experiment-tracking platform; it's a board for the human decisions around AI work, not a replacement for the tools that run the training.
The structure, and why each part is there
Every card starts as something you're actually unsure about, "would X help with Y", rather than a feature description. If the answer is already known, it isn't an experiment, it's just a task, and it belongs on a different board.
A low WIP limit here stops five vague experiments running at once with nobody's full attention on any of them. It also forces a decision: finish evaluating one before starting the next.
Every experiment passes through here before it goes anywhere else. The card records what was measured and what the result was, whether that result is good, bad, or genuinely inconclusive. Inconclusive is an allowed answer, not a reason to leave the card sitting untouched.
Both are real endpoints, not a success column and a failure column in disguise. A shelved experiment isn't wasted work, it's an answer to a question the team no longer has to wonder about.
What was tried, what the result was, and why it stopped: written down with page history, so it's findable when someone proposes the same idea again next quarter, instead of getting re-run from scratch. Six months on, that's usually the difference between a team that remembers it already tried something and one that quietly reruns it.
A training script or a feature-flagged rollout tied to a card moves with the branch and the merged pull request, so the state of the code and the state of the experiment don't drift apart from each other.
Experiments still take someone's time to design, run and read, and planning accounts for who's actually available that sprint. A team member on leave isn't a team member you can schedule three experiments against.
How to use it
- 01Start a workspace. It opens with a sample project already on the board, so the flow from hypothesis to shipped-or-shelved is visible before you add real work. Free for up to five people, permanently.
- 02Phrase every new card as a question, not a plan. If you can't state what you're unsure about, it probably isn't an experiment.
- 03Set the Experiment column's WIP limit low. Two or three running at once is usually enough for a small team to actually evaluate properly instead of half-watching five.
- 04Write the evaluation before deciding the outcome. Record the number first, decide shipped or shelved second, not the other way round. Deciding the outcome first and backfilling a rationale is how mediocre experiments quietly ship.
- 05Log every shelved idea in the wiki, briefly. A few sentences on what was tried and why it stopped is enough to save someone a wasted week later.
A team that ships one out of every four experiments and logs the other three properly is doing better work than a team that ships one out of every four and lets the other three disappear without a trace. The ratio isn't the problem. The silence is, and the shelved column with its wiki entry is what keeps the team's memory intact instead of resetting every quarter.
- Most experiments won't ship, build the board around that, not against it
- Shelved needs to be a real, visible outcome, not a silent deletion
- A logged reason an idea didn't work is worth more than the idea itself, next time it comes up
Common questions
No. ShipSprint doesn't integrate with model registries or experiment-tracking platforms. This board tracks the human decisions around AI work (what to try, what was found, what shipped) rather than metrics, model artifacts, or training runs themselves.
Because archiving makes it invisible, and invisible experiments get re-run. A shelved column keeps the outcome on the board where the next person planning similar work will actually see it, with the wiki entry explaining why it stopped attached.
ShipSprint connects to Claude and ChatGPT, so you can ask in plain language what's already been evaluated, shipped, or shelved, and get an answer drawn from the actual board and wiki history instead of relying on someone's memory.
That's a definition your team should set, not one the board imposes. Some teams call it shipped once it's live for any real users, others wait for a specific accuracy bar. Whatever the line is, write it down once so "shipped" means the same thing on every card.
Forecasts are calculated from measured velocity as sprints complete, and most teams get more reliable numbers by tracking hypothesis-to-experiment throughput separately from committed feature work, rather than folding an uncertain experiment into a delivery date it was never going to reliably hit.
Yes. All templates are included on the Free plan, which covers up to five users and two projects and doesn't expire. See pricing for larger teams.
Related pages
See it on your own work.
A workspace your whole company will actually use is 60 seconds away. No card, no risk, nothing to install.
14-day full-access trial · sample project included · no card required