The short version: a good DAM taxonomy is shallow, unambiguous, and backed by a controlled vocabulary so search terms don’t fragment. The discipline behind it isn’t guesswork — it’s the same machinery that runs the world’s medical and art databases: a preferred term for each concept, synonyms mapped to it, and broader/narrower relationships (the standard calls them BT and NT). This guide builds that at your scale, with real examples, and honest rules for what earns a place and how to govern it. If you’re still deciding categories vs keywords vs tags, start there first.
Why taxonomy makes or breaks findability
Every complaint people blame on their DAM — “I can never find anything,” “we keep re-shooting things we already have,” “nobody trusts the library” — is usually a taxonomy problem wearing a software costume. A taxonomy is the agreed structure that tells everyone where an asset lives and how it’s described, and when it’s good, findability feels effortless; when it’s absent, the library quietly fills with duplicates because it’s faster to make a new file than to find the old one. The reassuring part is that this is a solved problem — librarians and information scientists have spent a century working out how to make large collections findable, and the rules are written down. You don’t need a perfect system a cataloguer would admire; you need a structure your actual team can apply under deadline, built on principles that are known to work.
Designing the category tree
Categories are the hierarchy — the folders-that-aren’t-folders where an asset primarily lives. Two principles keep a tree usable. The first is mutual exclusivity within an axis. Information science calls the units of a good classification facets: in B.C. Vickery’s classic definition, facet analysis is “the sorting of terms… into homogeneous, mutually exclusive facets, each derived from single characteristic of division.” In plain terms, each level of your tree should split on one thing — by product, or by campaign, or by asset type — not a muddle of all three, so that for any asset there is one obvious home rather than three plausible ones. If your team routinely hesitates between two branches, that’s a design bug; merge or rename until the choice is obvious.
The second principle is keep it shallow, and here honesty matters more than a rule of thumb. There is no standard that prescribes a maximum depth — ANSI/NISO Z39.19, the US standard for controlled vocabularies, sets no limit, and the famous “seven, plus or minus two” figure people quote comes from a 1956 psychology paper on short-term memory, not from any taxonomy standard. So treat depth as a usability trade-off you own: every extra level is another decision the person filing has to get right, and a tree so deep that people give up and dump everything at the top is worse than a shallow one. Most working libraries land around three or four levels not because a standard says so, but because that’s where the filing stays reliable.
One subtlety worth knowing: a formal taxonomy actually permits a term to sit under more than one parent (Z39.19 defines a taxonomy as terms in “one or more parent/child… relationships”). That’s useful in a giant reference vocabulary, but for a browsable DAM it usually pays to give each asset one primary home in the tree and push everything that cuts across it — season, campaign, rights status, colour space — into keywords, which an asset can carry many of.
The prettiest taxonomies I’ve seen in a vendor demo are the ones that fail fastest. They’re deep, elegant and built by one person who holds the whole logic in their head — and then that person goes on holiday, three other people file the week’s shoot, and by month two the same kind of asset lives in four different branches. The taxonomies that survive are the boring, shallow ones where the filing decision is so obvious that a new hire gets it right on day one. Design for the tired Friday-afternoon uploader, not for the demo.
The controlled vocabulary — and the three relationships that make it work
If categories are where an asset lives, keywords are what it’s about — and they only work if everyone draws from the same approved list. That list is a controlled vocabulary, and the reason it works is that it encodes three specific relationships between terms. These aren’t our invention; they’re the backbone of the NISO standard and every serious vocabulary built on it:
- Equivalence — USE / UF. Many words mean the same thing, so you pick one preferred term and map the rest to it as synonyms. In the standard’s notation, the non-preferred word carries a
USEpointer and the preferred term carries the reciprocalUF(“used for”). This is what stops “car,” “cars” and “automobile” splitting your search three ways. - Hierarchy — BT / NT. Terms nest from general to specific with a broader term (
BT) and narrower term (NT): tagging something a convertible also makes it findable under its broader term, vehicles. - Association — RT. Some terms are related without being broader or narrower — a related term (
RT) cross-links them, the way a library links “Rugs” and “Carpets” so a search for one surfaces the other.
A fourth tool ties it together: the scope note (SN), a short line that pins down exactly what a term covers when its everyday meaning is fuzzy. Get these four things right — preferred terms, mapped synonyms, a shallow hierarchy, and scope notes for the ambiguous ones — and your keywords become a search system instead of a pile of labels.
The test I run on any vocabulary is to search for the wrong word on purpose. Type “sofa” when the library calls it “couch,” “car” when someone filed “automobile.” If nothing comes back, the vocabulary has no equivalence layer, and every user is one synonym away from concluding the asset doesn’t exist and re-making it. Mapping synonyms with USE/UF is the single highest-return hour of taxonomy work there is — it’s cheap to do and invisible until the day it saves someone an afternoon.
What a real controlled vocabulary looks like
This is easier to trust when you see it working at scale, so here are two you use the results of every day without knowing it. The first is MeSH, the Medical Subject Headings the US National Library of Medicine uses to index the world’s medical literature. When a doctor searches PubMed for “cancer,” they get everything — because MeSH maps a whole cloud of synonyms onto one preferred term:
| Element | Value |
|---|---|
| Preferred term (descriptor) | Neoplasms |
| Synonyms mapped to it (entry terms) | Cancer · Tumor · Tumors · Neoplasia · Malignancy |
| Position in the hierarchy | Tree number C04 (under Diseases) |
| Scope note | “New abnormal growth of tissue.” |
Everything in that record is a rule from the last section made concrete: one preferred term, a set of synonyms (MeSH calls them entry terms) folded into it, a place in a broader/narrower tree, and a scope note that says what it means. The Getty Art & Architecture Thesaurus — which catalogues the material world of art and architecture across roughly 74,000 concepts — shows the same machinery on a single object. Its record for cinnabar (mineral) carries the preferred descriptor, synonyms mapped in as “used-for” terms (cenobrium, natural vermilion), a scope note (“a soft, dense, red, native ore composed of mercuric sulfide…”), a broader-term chain (materials → inorganic material → mineral → cinnabar), and a related term pointing to the separate concept cinnabar (pigment). You are not going to build 74,000 concepts. But your DAM vocabulary should do exactly what these do at your scale — and any DAM worth its licence lets you record a preferred term, its synonyms and its parent, which is all this is.
How to decide what earns a place
The fastest way to ruin a vocabulary is to add terms speculatively — every keyword someone might want, until the list is so bloated no one can pick the right term. The standard has a cleaner discipline for this, called warrant: a term earns its place only when something justifies it. Z39.19 names three kinds, and all three are worth mining before you add anything:
- Literary warrant — the words that actually appear in your content and captions. If your photographers keep writing “golden hour,” that’s a candidate term; a word no one ever uses is not.
- User warrant — the words people actually search for. Your DAM’s own search logs are the best taxonomy research you have: the queries that return nothing are a to-do list of missing synonyms and terms.
- Organizational warrant — the terms your business specifically needs, whether or not anyone else would: product lines, campaign codes, region names, rights states.
Run a proposed term past those three and most bad additions disqualify themselves. It’s the difference between a vocabulary that grows because it’s needed and one that grows because it can.
The required metadata schema
Alongside the vocabulary sits the question of which fields every asset must carry. A common mistake is offering fifty optional fields that stay empty; the fix is to decide a small required set and enforce it at ingest, so an asset can’t enter the library half-described. Build that set on a standard so it stays portable. The two that matter for visual assets are Dublin Core — the 15-element model (Title, Creator, Subject, Description, Date, Rights and so on) maintained by the DCMI — and IPTC Photo Metadata for the photo-specific rights and description fields. A sensible required core:
| Field | Why it’s required |
|---|---|
| Description / Caption | What the asset shows, in plain language |
| Keywords | From your controlled vocabulary — the search layer |
| Creator | Who made it (photographer, designer, agency) |
| Copyright Notice & Rights Usage Terms | Who owns it and how it may be used |
| Status & licence-expiry date | Operational fields that stop an out-of-licence asset shipping |
One honest correction to a widespread belief: IPTC does not designate any field mandatory — in the specification every core field is optional (its cardinality starts at zero). “Required” is a decision you make and enforce in your DAM, not something the standard imposes. And because these fields are only as valuable as they are durable, confirm your DAM writes them back into the file as IPTC/XMP rather than holding them only in its own database; our metadata fidelity ranking shows how much that varies between tools.
Naming conventions
File names are metadata too, and a consistent convention pays off every time someone exports, shares or sorts a batch. You don’t have to invent the rules — research-data and archive teams have published them for years, and they agree. Harvard Medical School’s data-management guidance is a good, blunt example: no spaces (“many computer systems cannot handle spaces in file names” — use dashes or underscores); dates in ISO 8601 (YYYY-MM-DD or YYYYMMDD, which “makes sure all of your files stay in chronological order”); zero-padded sequence numbers (001, 002 … 010, so they sort correctly); short but meaningful names of roughly 40–50 characters using only letters, numbers, dashes and underscores; and document the convention in a README kept with the files. Read left-to-right from general to specific and a name explains itself:
- Good:
2026-05-04_spring-launch_hero_012.jpg— sorts by date, says what it is, survives every system. - Bad:
IMG_4821.jpg(says nothing) ·Final Photo (2)!.jpg(spaces, brackets, punctuation, and a “final” that never is).
Naming conventions feel like bikeshedding until you’ve tried to find one image across three years of a team that never agreed on one. Have the argument once, write the pattern down, and — this is the part people skip — apply it automatically on ingest so it doesn’t depend on a tired human at 5pm. The single highest-leverage rule is the ISO 8601 date: 2026-05-04 instead of May 4 or 5/4/26. It sorts chronologically for free, it’s unambiguous across countries, and it ends the annual argument about whether the month or the day comes first.
Governing it over time
A taxonomy is not a one-time project; it’s a living thing that drifts the moment you stop tending it. Here, too, the standard has a playbook worth copying. Z39.19’s guidance on maintenance is concrete: establish procedures for adding, modifying and deleting terms, and on every term record “note the date of each change and identify the individual responsible for it.” New terms enter as candidates nominated by the people indexing and searching, then get reviewed — ideally by a named editor or a small editorial board — against the warrant test above. And nothing is ever silently deleted: when a term is replaced, you record the change in a history note and leave a USE reference from the old term to the new one, so a search for the retired word still lands people in the right place.
Translate that to a DAM and it’s refreshingly small: give the taxonomy a named owner, put a recurring review on the calendar (quarterly is common) to fold in warranted new terms and retire dead ones, and treat changes like code changes — proposed, reviewed, documented, never improvised. The libraries that stay searchable for years aren’t the ones with the cleverest initial design; they’re the ones somebody kept weeding. Where this work most often surfaces is a migration, when a messy old scheme has to be mapped into a clean new one — far easier if you’ve been maintaining it all along.
Sources & references
- ANSI/NISO Z39.19 — Guidelines for the Construction of Controlled Vocabularies — the source of the USE/UF, BT/NT, RT relationships, scope notes, warrant and maintenance guidance cited here, accessed July 2026.
- U.S. National Library of Medicine — MeSH record types — the “Neoplasms” descriptor and its entry-term synonyms used as the worked example, accessed July 2026.
- Getty — About the Art & Architecture Thesaurus (AAT) — facets, preferred/used-for terms and the “cinnabar” record structure, accessed July 2026.
- Library of Congress — Subject Headings Manual, H 370 — real BT/NT and RT examples in a working vocabulary, accessed July 2026.
- Harvard Medical School — File Naming Conventions — the no-spaces, ISO 8601 date, zero-padding and length rules cited here, accessed July 2026.
- DCMI — Dublin Core Metadata Element Set — the 15-element schema model, accessed July 2026.
- IPTC Photo Metadata Standard — the description and rights fields, all specified as optional, accessed July 2026.
- ISO 25964 — Thesauri and interoperability — the international standard for thesaurus construction and interoperability, accessed July 2026.
- PhotoLib methodology — how we research and test DAM tools. See our methodology.
Keep reading
FAQ
What is a DAM taxonomy?
It is the agreed structure that makes a library findable: a category hierarchy assets are filed into, a controlled vocabulary of keywords used to describe them, and the metadata fields every asset must carry. The controlled-vocabulary part follows the same relationships used by professional vocabularies - a preferred term for each concept, synonyms mapped to it, and broader/narrower links - so that search returns everything on a topic instead of only the assets that happened to use one exact word.
What is the difference between broader/narrower terms and USE/UF?
They are two of the three standard relationships in a controlled vocabulary. Broader term (BT) and narrower term (NT) build the hierarchy - 'convertible' has the broader term 'vehicles', so tagging the convertible also makes it findable under vehicles. USE and UF handle synonyms: the non-preferred word 'automobile' carries a USE pointer to the preferred term 'car', and 'car' carries the reciprocal UF ('used for') 'automobile'. The third relationship, RT (related term), cross-links terms that are associated but neither broader nor narrower.
How deep should a DAM category tree be?
There is no standard that sets a maximum - ANSI/NISO Z39.19 prescribes none, and the popular 'seven plus or minus two' figure comes from a 1956 psychology paper on memory, not from any taxonomy standard. Treat depth as a usability trade-off you own: every extra level is another filing decision to get right, and most working libraries settle around three or four levels because that is where filing stays reliable. Keep each level mutually exclusive so any asset has one obvious home, and push anything that cuts across the tree into keywords.
How do I stop keywords fragmenting into near-duplicates?
With the equivalence layer of a controlled vocabulary: choose one preferred term per concept and map the synonyms to it, exactly as the medical MeSH vocabulary maps 'Cancer', 'Tumor' and 'Malignancy' onto the single descriptor 'Neoplasms' so a search for any of them returns everything. In practice that means maintaining a short synonym list (the standard's USE/UF references) and reconciling AI-suggested tags against your approved list rather than letting every user invent their own label.
Which metadata fields should be required for every asset?
Keep the required set small and enforce it at ingest. A sensible core built on the IPTC and Dublin Core standards is Description/Caption, Keywords (from your controlled vocabulary), Creator, and the rights fields Copyright Notice and Rights Usage Terms, plus two operational fields most workflows need - an approval status and a licence-expiry date. Note that IPTC itself makes no field mandatory; 'required' is a policy you set and enforce in your DAM. Whatever you choose, confirm the DAM writes those fields back into the file as IPTC/XMP so they stay portable.