- Home
- Research skills
- Research Repository and Knowledge Curation
Research Repository and Knowledge Curation
Structure your accumulated research so a question can be answered from it, with ageing findings flagged, contradictions visible and maintenance that actually happens.
Free. Works on Claude, ChatGPT, Gemini or any assistant that accepts a skill file.
What this skill does
The method, encoded.
Organisations rarely lose their research. The files are there, in folders named by year and agency. What is lost is the ability to answer a question from them, because nobody arrives at a repository looking for a document. They arrive with a question, and a list of study titles is not an answer.
This skill treats a repository as an answering system rather than a filing system. It starts from the questions people actually ask, sets the unit of curation as the finding rather than the report, and makes provenance mandatory: every entry carries its base, population, fieldwork period and confidence, because an entry that cannot be traced back to its study is a rumour with a search index.
It is strict on three things that decide whether a repository survives. A small controlled vocabulary beats a large uncontrolled one. Ageing evidence needs an explicit status (current, dated, superseded, retired), because a stale finding served as current does more damage than no repository. And two studies that disagree both stay discoverable, with the disagreement recorded rather than resolved on the searcher's behalf.
It also says the thing most repository projects learn too late: repositories fail on maintenance, not on design.
Best used for
- Making several years of studies interrogable rather than merely stored
- Stopping the same research being commissioned twice
- Preserving institutional memory when researchers leave
- Designing a research taxonomy that will still work at scale
- Preventing out-of-date findings being read as current
- Assembling and standardising a corpus before cross-study synthesis
- Diagnosing why an existing repository has stopped being used
Typical inputs
What you give it.
The real questions the repository must be able to answer, from a request log, Inventory of existing research with dates, methods, samples and file locations, A named curator with allocated time, Consent, contractual and confidentiality terms per study, Optional analysis outputs rather than only final reports, Optional record of the decisions each study informed, Optional quality review status per study, Optional search or usage analytics from an existing store
Typical outputs
What you get back.
Repository specification with question set and named acceptance tests, Study and finding record schemas with mandatory provenance fields, Faceted controlled vocabulary with a definition and boundary case per term, Status convention (current, dated, superseded, retired) with review and event triggers, Contradiction register with diagnosed causes and what would settle each, Access, redaction and retention model per class of material, Contribution workflow embedded in project close, with curator review, Quarterly use report covering retrievals, failed searches and overdue reviews
Method coverage
What the skill works through.
- What a research repository is actually for
- Designing from the questions people ask
- Why the unit of curation is the finding, not the report
- The entry schema and non-negotiable provenance
- Small controlled vocabulary versus large uncontrolled tagging
- Facets rather than a hierarchy
- The ageing problem and the four-status convention
- Review dates and event triggers
- Duplicates, corroboration and supersession
- When two studies contradict each other
- Access, confidentiality and participant data in a searchable system
- Contribution workflow and the curator's role
- Backfill: how far back to go
- Measuring use, and what a failed search tells you
Download
Free skill. One file.
Enter your email once. Every skill you download after that takes a single click.
How to install
Add the skill file and the five kernel protocols to a Claude Project, a ChatGPT Project, a Gemini Gem, or paste them at the top of any assistant conversation. Then give it your real research material, not a description of it.
Download skillQuestions
Common questions.
What should a research repository actually contain?
Findings, with studies as parent records. Nobody searches a repository for a report, so a store of documents pushes the work back onto the searcher. Curate individual claims that stand on their own, each carrying its base, population, fieldwork period, method and confidence, and link them to the study they came from. Extracting findings is the main cost of a repository and it is also the entire value.
How should research findings be tagged?
Use a small number of independent facets rather than one deep hierarchy, because a finding is about a topic, a population, a market and a decision at the same time. Keep the vocabulary small and controlled: aim for roughly fifteen to thirty topic terms, write a definition for each including the boundary that separates it from its neighbour, and require the curator's approval to add a new one. A small controlled vocabulary beats a large uncontrolled one every time.
How do you stop a repository serving out-of-date findings?
Give every finding a status rather than every study, since findings inside one study age at different rates. Four statuses work: current, dated, superseded (with a link to what replaced it) and retired. Set review periods by volatility, add event triggers such as a product, pricing, policy or market change, and make a missed review downgrade the entry to dated automatically rather than leaving it current by default.
What do you do when two studies in the repository contradict each other?
Keep both, keep both discoverable, and link them with a record that states what each found, the diagnosed cause of the difference (different population, wording, period, method, or genuine disagreement), which is better evidenced and why, and what would settle it. Showing only the more recent or more convenient finding makes an analytical judgement invisibly on behalf of everyone who will ever search. The contradiction record is frequently the most useful entry in a repository.
Why do research repositories fail?
Almost always on maintenance rather than design. The pattern is predictable: contribution falls once the founding projects close, tags drift because nobody reviews them, a stale finding is served as current, trust breaks, and people go back to asking the longest-serving researcher. The fix is a curator with allocated hours and contribution written into project close, not a better schema.
How far back should we curate our existing research?
Only as far as people are actually asking. Comprehensive backfill is where repository projects consume their budget and never launch the forward workflow. Curate forward from today, and backfill the specific studies that appear in your request log, which is usually ten to twenty rather than everything.
Can old research be deleted once it is superseded?
No. Keep superseded and retired entries, linked to whatever replaced them. The historical record is the only thing that makes change over time visible, and an organisation that deletes its old answers ends up believing it has always known what it currently knows.
What about participant data and confidentiality in a searchable repository?
Treat searchability as a new purpose that consent given for a report does not automatically cover. Decide per class of material what may be entered, what is entered but access-restricted, what is redacted on entry, and what retention period applies. A verbatim that was safe in an aggregated report can identify someone once it is tagged by market and segment and retrievable organisation-wide. Restricted entries should be visible as restricted, so a searcher knows material exists rather than concluding nothing was found.
Do we need a repository at all?
Below roughly fifteen studies, probably not. A single well-maintained index listing study, date, question, sample, method, headline findings and file location will outperform a repository, cost almost nothing and need no curator. And if nobody has time allocated to maintain it, do not build one at all: an unmaintained repository is worse than none, because it looks authoritative while serving stale evidence.
The skill chain
Works well with.
Research where people already are.
Analyse it where you already work.
Yazi helps researchers conduct surveys, AI interviews and longitudinal research directly through WhatsApp.
%202.png)
