What spend analysis is actually for
Spend analysis exists to make four things possible: a procurement target you can defend, the mandate to go after it, a road map for delivering it, and the compliance that makes the savings stick. It is a means, not an output. What the data has to contain depends on which of those you are after.
The mechanics are unglamorous. Turn transaction data from your finance and procurement systems into a single, deduplicated, classified view of what the organisation buys, from whom, and for which part of the business. A first pass takes two to four weeks.
Before any of that, check one number: what share of spend actually goes through a purchase order. Where an organisation runs a purchase order system with genuine adoption, that is the best foundation available, better than anything reconstructed from invoices afterwards, because the detail is captured at the point of ordering rather than inferred later. Where PO coverage is thin, you are working from accounts payable, the ledger and the invoice documents themselves. The two routes cost different amounts and reach different places, so establish which one you are on first.
Most published guides stop there, at extract, cleanse, classify, analyse. That is the sequence, and it is not the point. The question that comes first is what the analysis is for, because that decides how far you have to go and when you are allowed to stop.
Spend transparency is the precondition, not the project
You cannot write a procurement strategy without a view of the spend. There is no version of that sentence that is controversial, and yet a surprising number of mid-sized and smaller companies do not have basic spend transparency at all. Not imperfect transparency. None.
Larger organisations usually have something, and many have a genuinely good spend cube. That is not the problem. The problem is that a good spend cube is built for some questions and not others, and having one is not the same as having transparency at the level the question in front of you requires. That distinction is what the rest of this article is about.
Without it, everything downstream is guesswork. You cannot say which categories matter, so you cannot prioritise. You cannot set a target you are willing to defend. You cannot segment the supply base, so supplier relationship management has nothing to stand on. You cannot tell whether a saving happened. Modern procurement work is not difficult without spend transparency. It is not available.
What you are going to use it for, and how granular it has to be
This is the decision that determines cost, duration and whether the exercise succeeds. Most published guidance treats spend analysis as one thing with one output. It is not. There are at least four distinct uses, and they have different data requirements.
- A procurement strategy, a target and a mandate. Where the money goes, where the opportunity sits, what the ambition should be, and in what order to attack it. This is also how procurement gets a mandate. “Procurement should do more” is not a proposition a board can approve. “There is NOK 1.4bn of addressable third-party spend, here is what we believe is available on it, here is the sequence and here is what we need to go after it” is. The number is what converts an aspiration into a decision, and you cannot produce the number without the analysis.
This is the use case most first analyses serve, and it is the least demanding on data. Supplier, category and business unit across twenty-four months is enough. Good enough is genuinely good enough here.
- Category strategies for individual categories. The input to deciding what happens to a single category over the next three to five years. Needs everything above as the frame, plus depth inside the category: volumes, price points, contract terms and expiry dates, and the specification wherever it is recorded.
Be clear about where that depth comes from, because this is where expectations of a spend tool usually break. It does not come out of the spend cube. It comes from a separate collection exercise: data requests to the suppliers themselves, a contract review, and volume and specification data held by the business rather than by finance. The same applies to running an actual sourcing event, where the cube will not carry the line-level detail a market process needs to price against.
-
Cross-cutting commercial levers in isolation. Payment terms are the clearest example. Moving terms across every supplier at once is a working capital exercise that ignores category boundaries entirely. Index clauses, currency exposure and consolidation of duplicated suppliers behave the same way. What you need is spend joined to contract terms, not more detail per transaction.
-
Compliance and price assurance. Are we actually buying from the supplier we selected, and are we being invoiced the price we agreed? This is a different order of requirement. It needs line-item data, joined to the contract, the agreed price list or the agreed pricing principles, and refreshed continuously rather than annually. An analysis built for use case 1 cannot answer it, and no amount of reworking will make it.
| Use case | What it answers | Minimum data | Where it comes from | Refresh |
|---|---|---|---|---|
| Procurement strategy | Where is the money, where is the opportunity, what is the ambition | Supplier, category, business unit, 24 months | The spend cube | Annually |
| Category strategy | What do we buy here, from whom, at what price, on what terms | The above, plus volumes, prices, terms and expiry within the category | The cube for the frame, then supplier data requests, contract review and business data | At the start of each category |
| Cross-cutting commercial levers | Where can we move payment terms, index clauses, currency exposure | Spend joined to contract terms | The cube joined to the contract register | Annually, or on renewal |
| Compliance and price assurance | Are we using the chosen supplier, and paying the agreed price | Line item, joined to contracts and agreed prices or pricing principles | The cube at line-item level, joined to agreed price bases | Monthly or continuous |
Read the fourth column before the third. A spend cube is a planning and compliance instrument. It is not an execution instrument. It exists to give you good enough data to decide where to act and to check afterwards whether what you agreed is actually happening. It will not run a category strategy or a sourcing event on its own, and organisations that expect it to are disappointed by a tool that is working exactly as designed.
That boundary is a statement about today rather than a permanent one. A cube fed with invoice line items and joined to contracts and to the wider business data would begin to carry execution work, and that is the direction of travel. It is not what we see in practice now, and getting data structured to that standard takes years, particularly in large groups with several system estates. Plan on the boundary holding for the next five to ten years, and treat anything better as an advantage you have earned rather than a capability you can assume.
Granularity is set by the use case, not by ambition. Two failure modes follow from ignoring that. The first is chasing line-item detail for a question that only needed supplier and category, which adds months and cost for no decision that changes. The second is promising compliance monitoring on a dataset that never carried line items, which fails the first time someone asks for the evidence.
Decide the use case, then decide the data. Not the reverse.
Institutionalise it, and do not build it yourself
A one-off analysis decays from the day it is delivered. The value is in having it standing, refreshed, and available when a question arrives. That means a tool, whether bought or built.
Our advice is almost always to buy. Building your own ends up more expensive than the licence, less flexible when you need functionality you did not anticipate, and less capable than products that have been developed against hundreds of implementations. The cost of building is never the first version. It is the second and the third, and the person who has to maintain it after the person who wrote it leaves.
The tooling has also moved. Products now exist that read invoice documents directly and extract line-item detail where the invoice carries it, which puts use case 4 within reach of organisations that could not previously get near it. That capability is worth understanding before you scope anything, because it changes what is possible. It does not create detail that the invoice never contained.
The one legitimate case for building is a data model genuinely unlike anything on the market. It is rarer than the people proposing it believe.
Step 1. Write down the question
A spend analysis built to answer “where does our money go” produces a dashboard. One built to answer “which three categories do we work on in the next twelve months, and what is each worth” produces a decision.
Write it down first, and write down which of the four use cases you are serving. Everything below follows from it.
Step 2. Pull the data, and know what each source can and cannot tell you
| Source | What you get | The limitation |
|---|---|---|
| Purchase orders | Line-item detail, quantities, unit prices | Covers only the spend that went through a PO |
| Accounts payable and invoices | Near-complete coverage of what was paid | Often header level only in the ERP, even where the invoice document itself carries lines |
| General ledger | Complete, reconciles to the accounts | A supplier name and an account string, nothing more |
| Invoice documents and their attachments | Line-item detail where the supplier provided it, including detail carried in appendices such as specifications and timesheets | A separate data pull from the document archive, not the ERP extract; nothing to extract if the invoice was issued as a single line |
| Contract register | Terms, expiry dates, owners | Usually incomplete and out of date |
Twenty-four months is the practical default. Twelve hides seasonality and gives no trend; thirty-six adds effort and pulls in an organisation that no longer exists.
Start at the top of that table, not the middle. Purchase order data is structurally the best source you can have: line item, quantity, unit price, requester, cost centre and category, all captured at the moment of ordering by someone who knew what they were buying. Everything below it in the table is an attempt to recover after the fact what the purchase order would have recorded for free. Where adoption is high, the analysis is faster, cheaper and better, and improving PO compliance is often a better investment than any tool.
The fourth row is the change worth noting where PO coverage is not there. A large share of indirect spend has historically been unreachable because the ERP holds only an invoice header, so classification rests on a supplier name and an account code. Extraction tooling reads the underlying invoice document and recovers the lines where they exist, which materially widens what is reachable and is what puts the compliance use case within range.
Two conditions apply, and the second is the one that disappoints people.
First, you have to source the documents themselves: the actual invoice files and their attachments, pulled from the document archive rather than the ERP extract, because much of the useful detail sits in an appended specification, parts list or timesheet rather than on the invoice face. Plan that as a separate data request with its own owner.
Second, extraction recovers what is on the document and cannot create what is not. Where the supplier invoices “consultancy services, one line, one amount”, there is nothing to find. The harder case is complex products and solutions, where the invoice may carry a line and a price and still tell you nothing about specification, configuration, scope or what actually drove the cost. Those are precisely the categories where you most want the detail, and they are the ones extraction helps least with. For those, the answer is not a better tool. It is asking the supplier.
Step 3. Strip out what you cannot act on
This is the step that separates a usable number from an impressive one, and we have not found it set out in any of the widely read guides. Before anything else, remove:
- Intercompany and internal recharges. In any Nordic group with a shared-services or holding structure this is routinely the single largest distortion of the top-line figure.
- Payroll and personnel costs.
- Taxes, duties, customs and statutory fees.
- Regulated tariffs where no supplier choice exists.
- Rent under a long lease, and financing costs.
- Grants, transfers and pass-through items.
What remains is third-party spend. Of that, addressable spend is the portion you can realistically act on within the horizon you care about. Say out loud which exclusions you have applied. “Our spend is NOK 800m” means nothing until someone asks what is in it.
Step 4. Fix the supplier master before you classify anything
Published guidance treats this as a spelling problem: group IBM, IBM Corp. and I.B.M. together. That is the easy half.
The hard half is legal entity structure. Match on the organisation number, not the name: organisasjonsnummer in Norway, organisationsnummer in Sweden, CVR in Denmark, y-tunnus in Finland. Names change, get abbreviated and get typed differently. Registration numbers do not.
Then decide, explicitly, how far up the corporate hierarchy you roll up, because the answer changes your number and your strategy:
- Roll up to the ultimate parent and your apparent leverage grows. Four subsidiaries of the same group become one supplier relationship worth four times as much.
- Do not roll up and you keep the view that matches reality, because the entity you contract with, the entity that invoices you and the entity that can actually agree a price are often the subsidiary, not the group.
Produce both views. Use the parent view to decide where to concentrate attention and the entity view to plan the negotiation. Presenting only the parent view creates a leverage story you cannot execute.
Step 5. Settle the currency convention
Any Nordic group buying across NOK, SEK, DKK and EUR has to choose a rate convention, and the choice silently creates or destroys apparent savings.
Three options: transaction-date spot, monthly average, or a fixed budget rate. There is no single correct answer. There is a wrong practice, which is using different conventions in different years and reporting the difference as a result. Pick one, apply it across the whole period, and state it on the front page of the output.
Step 6. Classify to a taxonomy you own
Two decisions here.
Which taxonomy. UNSPSC, created in 1998 by UNDP and Dun & Bradstreet, carries roughly 158,000 codes across four levels and works as a coverage check. ECLASS, published by the German ECLASS association, holds around 50,000 product classes and 23,000 properties in its 16.0 release of November 2025, and handles technical and direct materials better. CPV, established by Regulation (EC) No 2195/2002, is mandatory for public procurement notices across the EU and EEA.
None of them was built for procurement analysis. Use them as reference, then build your own structure around your supply markets, three to four levels deep, which is the conventional depth for a spend taxonomy.
How accurate is good enough. Vendors quote classification coverage above 95%. Coverage means a line received a code. It does not mean the code is right. The technical documentation behind those marketing figures is more candid: fully automated categorisation typically reaches 70% to 85% accuracy without a validation layer, and 80% coverage is the usual threshold for reliable analysis. Audit a random sample by hand and report measured accuracy alongside coverage.
Nordic-language supplier names and ledger descriptions also break classifiers trained on English, so budget for manual review of the top suppliers by value whatever the tool reports.
Step 7. Build the cube, then ask three questions
The standard output is a spend cube with three dimensions: supplier, category, and business unit or cost centre. That is the artefact. These are the questions that make it useful.
- Where is spend fragmented that should not be? The same category bought from many suppliers, at different prices, on different specifications, across sites.
- Where is spend concentrated that should not be? A single supplier holding a category with no contract, no benchmark and no alternative qualified.
- What is not under contract at all? Recurring spend with no agreement behind it is usually the fastest thing to act on and the easiest to prove.
Step 8. Know when to stop
We have not found a published stop rule anywhere, so here is a workable one. It is deliberately tied to the use case.
For a procurement strategy, stop when classification coverage is above 80%, measured accuracy on a hand-audited sample is above 85% across the top 80% of spend by value, exclusions are documented, and the top twenty suppliers are correct at both entity and parent level. Further precision rarely changes which three categories you pick.
For a category strategy, the estate-level thresholds still apply, but stopping is the wrong frame for the category itself. The numbers a negotiation runs off will not reach the required standard by cleaning the cube further. At that point you stop working on the cube and start collecting: supplier data requests, contract review, volume and specification data from the business. Treat that as a separate exercise with its own timeline, not as the last ten per cent of this one.
For compliance and price assurance, coverage thresholds are the wrong measure altogether. What matters is whether every contracted supplier is joined to its agreed price basis. A programme that covers sixty per cent of contracts completely is more useful than one that covers all of them partially.
What it costs, and when to abort
A first-pass analysis on a mid-sized Nordic group is typically two to four weeks of effort. Running past eight weeks almost always means the problem is the supplier master or an ERP estate left over from acquisitions. At that point it is a data project, not an analysis, and it should be renamed and rescoped rather than pushed.
Abort criteria are worth setting in advance. If more than a third of spend arrives with no supplier identifier that can be matched to a registration number, fix the master data first. Analysis on that foundation produces a number that will not survive its first challenge from finance.
Frequently asked questions
Why do a spend analysis at all?
Because every downstream decision depends on it. Setting a procurement ambition, prioritising categories, building a category strategy, moving payment terms across the supply base, segmenting suppliers, and proving a saving all require a view of what you buy and from whom. Without it, procurement is running on anecdote.
How granular does the data need to be?
That depends on the use case, and it is the most consequential question in the whole exercise. A procurement strategy works on supplier, category and business unit. Compliance and price assurance need line-item data joined to agreed prices. Building the second when you needed the first wastes months; promising the first can deliver the second destroys credibility.
Should we build our own spend solution or buy one?
Buy, in almost every case. Building costs more than the licence once maintenance is counted, adapts badly when you need new functionality, and rarely matches products refined across many implementations. The exception is a genuinely unusual data model, which is less common than its proponents claim.
How far back should we pull data?
Twenty-four months is the practical default. It captures seasonality and gives one year-on-year comparison. Go to thirty-six months only where contract cycles are long or you need to evidence a price trend.
What is the difference between spend analysis and spend analytics?
Spend analysis is the exercise: one classified, deduplicated view built to answer a defined question. Spend analytics is the standing capability, refreshed and used to monitor compliance and price movement. Do the analysis first. The capability is worth building only once someone acts on the output.
How much of our spend is addressable?
We work on 70% to 80% of third-party spend as a rule of thumb for most private-sector organisations, after intercompany, statutory and regulated items are removed. That is our working assumption rather than a published benchmark, and no published benchmark we have found states a methodology. Treat any figure above 80% with suspicion until the exclusion list has been checked line by line.
Who should own the output?
Finance should agree the baseline and the exclusions. Procurement should own the taxonomy and the classification. If those two do not agree on the number before it is presented, the number will be argued about instead of acted on.
Sources
- UNSPSC, created 1998 by UNDP and Dun & Bradstreet; GS1 US appointed code manager 2003.
- ECLASS e.V., ECLASS release 16.0, November 2025: approximately 50,000 product classes and 23,000 properties.
- Regulation (EC) No 2195/2002 on the Common Procurement Vocabulary (CPV), 5 November 2002.
- Sievo, vendor technical guidance on spend categorisation: 70% to 85% accuracy without validation layers; 80% coverage threshold. No methodology published.
- McKinsey & Company, “The role of spend analytics in the next normal”, August 2020: external spend commonly 40% to 80% of total cost.