Sign in Book a demo

Nobody Audits Their DAM. We Do It Before Every Migration.

Before a single file moves, we audit. Here’s what we find in every DAM library — and why clean data decides what AI can do with your brand.

Nobody Audits Their DAM. We Do It Before Every Migration.

On a recent migration, we touched 99% of the customer’s files before a single one landed in Collage. The library wasn’t in bad shape. That’s just what a real audit turns up.

We’ve done enough of these now that the findings have stopped surprising us. The same four problems show up across libraries of every size, in every industry, built on every platform. Left alone, they quietly tax the team every week. They also put a ceiling on what AI can do with a brand library. The teams investing in clean content data right now are pulling ahead in a race most of the market doesn’t realize has started.

Here’s what they are, why they compound, and what becomes possible when you actually fix them.

How the Audit Works

Before we migrate a single file, we run an audit. It draws on three inputs.

The first is the library itself. Our AI reads the signals already embedded in it: filenames, tags, folder structures, custom field patterns, and much more. Most libraries have effective organizational logic buried inside them; it’s just inconsistent, accumulated over years of different contributors and workflows.

The second is discovery. Before the audit runs, we sit down with the team to understand how they actually work — their processes, priorities, distribution channels, primary and secondary audiences, and the workflows the library needs to support. That context shapes what “good” looks like for their specific business, not a generic best practice.

The third, when it’s available, is supplementary data the customer provides directly: product information, channel mappings, brand hierarchies, anything that enriches what the library can tell us on its own.

Together, those inputs let us surface what’s actually in the library versus what teams assume is there, and reorganize it into something coherent. The process takes days, not weeks. What it reveals is almost always the same.

Finding 1: The Folder Structure No Longer Matches How Teams Work

Libraries accumulate containers. A folder gets created for a campaign. Then another for a product line that later got renamed. Then a workaround folder because nobody could find the original. Then a subfolder that made sense to one person on one day and hasn’t made sense to anyone since.

Over time, the container structure stops reflecting how the team actually works. It reflects history.

We’ve seen container counts drop by as much as 70% in an audit. The content isn’t going away; the architecture around it gets rebuilt to match current workflows rather than past ones. Navigation improves not because there’s less content, but because the structure finally matches the mental model of the people using it.

Finding 2: Metadata Fields Exist. They’re Almost Empty.

Most DAM platforms make it straightforward to configure custom metadata fields. Most teams do configure them, usually carefully, at the time of setup. Then the day-to-day reality of content production takes over, and actually populating those fields slides into the category of things that should happen but don’t.

What we consistently find: thoughtfully designed field schemas sitting at single-digit coverage rates, sometimes years after initial configuration.

The signal needed to populate those fields is already in the library. Tags, filenames, and folder placement all carry implicit metadata. When customers can provide additional context — product feeds, channel data, brand hierarchies — we factor that in too. Our AI works from those combined inputs to fill the gaps, and we’ve seen coverage move from near-zero to 90% or higher on individual fields. The underlying information was always there. The audit pulls it into structure.

Finding 3: Tags Have Become Noise

Tags are usually the first thing to degrade. The problem is rarely volume on its own. It’s duplication and drift. “Footwear,” “shoes,” and “Shoes” become three tags for the same concept. A new platform integration introduces auto-generated format tags that overlap with what already exists. Casing drifts between contributors. Synonyms accumulate. A new team member develops slightly different tagging habits than the people who came before them.

The result is a taxonomy where searching for one tag misses the assets filed under its three near-identical siblings. Findability degrades not because tags are missing, but because the same concept is fractured across multiple labels.

The audit consolidates duplicates and synonyms into a single canonical tag, and we’ve seen total tag counts drop by more than half in the process. Findability improves because the tags that remain are actually doing navigational work, and a leaner, more consistent tag set is easier to apply correctly going forward — which is how you prevent the same drift from happening again.

Finding 4: Filenames Are Effectively Unsearchable

Filenames carry more inferential weight than most teams realize. A well-structured filename tells you what an asset is, what campaign or product it belongs to, which version it represents, and sometimes where it should live in the taxonomy, all without opening the file. For AI-assisted workflows, filenames are often the primary signal used to infer metadata at scale.

What we usually find instead: mixed conventions across teams and time periods, generic names like “Hero Image Final,” copy artifacts from version management (“Copy of,” “(1),” “FINAL_v2”), and no consistent logic from one folder to the next. When filenames are inconsistent, that inferential signal breaks down, and every downstream process that depends on it gets harder to automate.

Why This Compounds

Each of these problems is tolerable on its own. Together they create a library that quietly costs the team time every week.

When findability degrades, teams stop searching and start recreating. Creative capacity and budget go toward producing assets that already exist somewhere in the library.

When metadata coverage is low, programmatic distribution becomes difficult or impossible. Filtering and syndicating content by SKU, brand, channel, or region requires someone to already know where everything lives and manually pull it.

When storage fills with duplicates, outdated versions, and orphaned files, you pay to store content that serves no one. More importantly, the signal-to-noise ratio for the entire library keeps declining.

The AI Layer Raises the Stakes

This is where the cost of accumulated debt stops being a productivity tax and starts being a strategic constraint.

AI tools are only as useful as the data they’re working with. Consistent naming conventions, structured metadata, and clean taxonomies aren’t organizational preferences; they’re the foundation that AI needs to do anything meaningful with a brand library. Teams building toward AI-powered content workflows need this infrastructure in place first. Without it, AI tools surface the wrong assets, make poor inferences, and require heavy human intervention to produce anything reliable.

A clean data layer isn’t the nice-to-have. It’s the prerequisite.

What the Audit Makes Possible

Call it a data upgrade more than a migration.

Back to that customer we mentioned at the top. Their library wasn’t in particularly rough shape by the standards of what we usually see. But by the time the audit was complete and we began migrating files into Collage, we had touched 99% of them. Tags normalized across 98.4% of assets. Filenames standardized on 70.4%. Storage locations remapped to the new org structure on 70.2%. Custom fields enriched or added on 42.8%.

The work that produced those numbers is the kind that never makes it onto a roadmap: inconsistent naming, bloated tag taxonomies, metadata that accumulated haphazardly over years of normal operations. We handle all of it in the migration window, before anything goes live in Collage, because the migration window is the one opportunity to restructure a library without disrupting the people actively using it.

Most teams carry that accumulated debt into their new platform and spend years working around it. You don’t have to.

Where This Is Heading

Right now, the audit is a discrete exercise we run before a migration. A concentrated effort to give teams a clean foundation before they start building on top of it.

The broader direction is toward making this kind of structural intelligence continuous rather than one-time. The teams getting the most from their DAM are already treating their asset library as a structured data layer, something that feeds downstream systems and workflows, rather than a well-organized folder of files. That framing is becoming more relevant, not less, as AI tools become more embedded in how brands produce and distribute content.

If you want to know what’s actually inside your library, we can show you.

Want to know what’s actually inside your library? Request an audit and we’ll show you what’s really there.

Get more value
from your content.

Book a demo