Essay · Workflow

Why My AI Prompt Library Died (And What I Keep Instead)

I built a careful personal prompt library. It still died. What travels now is role, residue, and a short ticket.

Graveyard of prompt notebooks labeled FINAL_v3, GOOD ONE??, and LOST, with a closed library binder in the fog
The good wording disappears. The pile does not. What I keep now is smaller and more portable.

I needed the same rewrite job for the third time in two weeks.

Not a fancy job. Make a flat section specific. Kill the productivity opening. Sound less like every other post about "using AI better".

I could remember the feel of the prompt that had worked last time. Short. Mean about adverbs. Told the model to start mid-scene.

I could not remember the actual wording.

So I spent fifteen minutes reconstructing it. Got something usable. Closed the tab.

A week later I did it again.

That was the quiet cost I had been ignoring. The good versions disappeared into chat history. The average ones got rewritten forever. I was paying a reconstruction tax on work I had already solved.

I am not special here. r/PromptEngineering is full of people who "kept losing the good prompts" into chat burial. Another thread is basically have you built a library, and did anyone still open it.

So I did what every productivity post told me to do.

I built a prompt library.

I actually used the rules that are supposed to make libraries work

I did not dump every experiment into a Notion graveyard.

I only saved prompts that had already produced something I would ship with light edits. Experiments stayed out until they earned a slot. That is the same selectivity you see in the "library that actually gets used" genre, like Hamza Aziz's write-up: prove it first, name it clearly, replace in place, organize around real work.

Every entry got a boring name and a one-line "when to use":

When a better version appeared, it replaced the old one. No v3 / v4 / final_FINAL pile.

I organized around real jobs, not abstract taxonomies. Writing. Editing. Outlining. Analysis. Planning. The folders matched Tuesday, not a framework diagram.

I also skipped the marketplace detour. Thirty-day "I tested five prompt libraries" pieces exist if you want one (this Design Bootcamp roundup is typical). Buying someone else's paragraph never fixed my voice problem, so I stayed with a personal list.

For a while it worked.

Opening a blank chat stopped feeling like inventing fire. I pulled a tested starting point, swapped in today's details, ran it. Output was more consistent. I felt briefly smug, which is usually a warning sign.

Then the library stopped getting opened.

Not because I got lazy. Because the library was solving the wrong missing piece.

Why a "good" library still becomes a graveyard

Two failure modes show up in every Reddit thread about prompt libraries.

Saved too early. You file the clever draft before it has survived real work. Later you open it, remember why you abandoned it, and stop trusting the whole list.

Find friction beats rewrite friction. If hunting the right entry takes longer than typing a rough new prompt, you stop hunting. The collection becomes a museum.

I hit both. They were not the main problem.

The third failure mode is the one that killed mine.

Three-panel comic: saved too early, too hard to find, no home for rules
I fixed the first two. The third one still ate the system.

Same prompt. Different week. Still wrong.

The library entry was stable. The job was not.

Last week's rewrite needed a US enterprise tone and a banned list of internal jargon. This week's rewrite needed founder-casual and a product name glossary. Same "rewrite for clarity" string. Different constraints. The prompt had nowhere durable to put those constraints, so I kept stuffing them into the chat like temporary sticky notes.

The model did what models do with sticky notes. It half-remembered them for one thread and forgot them the next.

Vendor features try to paper over this. ChatGPT Memory and the Memory FAQ are real product surface area. Claude Projects let you hang files and project-level instructions on a workspace. Useful. Still not the same as your standards living in plain files you control, especially once you bounce between products.

LangChain's note on context engineering is the cleaner label for what I was missing: the job is not a cleverer sentence in the prompt box. It is putting the right residue in the window when the work starts.

"Sound like me" is not a specification

Half my library said some version of write in my voice.

"Me" was not defined anywhere the model could reliably load. It lived in my head, in three old emails I liked, and in a soft rejection of LinkedIn cadence I could recognize but not paste.

So every run of "sound like me" was a coin flip dressed up as a system.

Multi-model made the lie louder

I do not stay in one chatbot. For anything that matters I want a second pass under a different engine. Same job, different judgment.

That habit wrecked the library faster.

I would paste a treasured prompt into one model on Monday and another on Thursday. Same text. Different house style bleeding through. I would "fix" the library entry to please whichever model I had open that day. The entry became a compromise that made nobody sound right.

The library was not a source of truth. It was a shared lie I kept editing.

What I keep instead: roles, not magic strings

I stopped trying to store the whole brain in a prompt paragraph.

I split the durable parts into three layers.

Role. Who this assistant is for this job. How it fails. What it must read before it answers. Failure modes matter as much as tone. "Never pad to sound smart" is more useful than "be concise".

Project residue. The files the role is allowed to lean on. A thin house-voice note. A glossary. Last good sample. A decision log. The messy middle of this project, not a universal mega-prompt.

Prompt. A short job ticket on top. "Rewrite section 3 for the VP audience. Keep the metric. Cut the origin story." Not the biography of my career.

Person wheeling glowing boxes labeled ROLE, RESIDUE, and TICKET past a dead filing cabinet of prompt printouts
Role and residue survive reuse. The ticket can stay short and throwaway.

The unit that survived reuse was the role plus the residue. The prompt became disposable again, which is fine, because disposable is honest for one-off instructions.

Old unit What I keep now Lives where
Rewrite prompt v7 Editor role + house voice notes Role instructions + one style file
Weekly report mega-prompt Comms role + meeting-notes cleanup rules Role + short trigger + sample outputs
"Think like a PM" blob PM role + open PRD / roadmap refs Role + project files

I do not need twenty roles. I need a few I actually reopen.

For day-to-day writing work I keep something like a comms/editor seat and an analysis seat. They share the same project files when the work is one product. They do not share a personality. The editor is allowed to be mean about adverbs. The analyst is not allowed to invent a metric to sound decisive.

When I want a second opinion, I change the engine under the same role. I do not rename the role after the model. The role is the asset. The model is the engine. If you bind "Alex = that one chatbot brand," you will rebuild Alex every time the vendor moves the floor.

I wrote the longer version of that setup as an employee handbook for an AI team. This piece is the narrower lesson: the handbook only helps if you stop pretending a prompt list is the handbook.

(If you want the org-side cousin of "stop worshipping libraries," there is also the harsher take that prompt libraries mostly teach facts, not situated judgment. I am not building a training program. I am just trying to stop reconstructing the same brief every Tuesday.)

Stop restarting. Edit the weak paragraph.

There is a smaller habit that pairs with roles and killed another time sink.

I used to trash a whole draft the moment the model went generic. New chat. New prompt. New lucky roll.

Now, if the bones are fine, I leave the role in place and point at the damage:

Split panel: smashing a typewriter crossed out versus a surgeon editing only the weak bit
Correct the weak paragraph. Do not re-interview a stranger.

Paragraph three is flat. Replace the example with something concrete from this week. Cut any adverb that is not doing work.

Rewrite the opening line only. It sounds like every productivity post on the internet. Start mid-scene.

Keep my structure. Tighten the middle. Do not add a summary section I did not ask for.

Surgical edits against a stable role beat re-pasting library entry #14 into a cold thread. You are correcting a colleague who already has the brief, not interviewing a stranger every time.

When a plain prompt is enough

I am not arguing that everything must be industrial.

Travel email. One-off brainstorm. Subject lines you will never reuse. A tight prompt in a blank chat is fine.

If you will not repeat the job inside thirty days, do not invent a role for it. Libraries die partly because people file everything. Selectivity is not a personality trait. It is how you keep the system small enough to reopen.

Setup cost only pays off on work that shows up weekly: status writing, draft cleanup, research synthesis, decision memos, the same three client formats.

What still belongs to you when the floor moves

Chat products change. Models get renamed, retired, rerouted. Prompt marketplaces will still sell you somebody else's paragraph.

OpenAI has been explicit about retiring older ChatGPT models and keeps shipping model release notes as the floor moves under daily habits. That is not a scandal. It is the product calendar. It is also why a prompt folder that only makes sense inside one vendor chat is a weak inheritance.

The question I care about now is smaller and more practical.

If the vendor moves the floor next quarter, what still travels with me?

A folder of prompts tied to one chat product travels poorly. Roles plus a few plain files travel better. The short job ticket on top can stay disposable.

I run that setup today in HaloMate, mostly because I can keep the same role and project files while swapping engines for a second pass without rebuilding the brief. The important part is not the logo. It is refusing to store my standards in a place that evaporates when the tab closes.

I still write rough prompts. I still throw some away.

I just stopped calling the pile a system.

Also on Medium / Write A Catalyst.