A feature store is shared infrastructure that computes feature values once and serves them to both training and inference from the same definitions. Enterprise ML needs one when the same features are reused across models, when training and serving compute them differently, or when inference must read features in milliseconds.
Writing about feature stores splits into two camps. Vendor pages explaining why every serious ML team needs one, and conference talks from companies operating at a scale you are probably not operating at. Neither is much help when you are the architect who has to decide whether this is infrastructure worth standing up this year, against everything else competing for the same budget.
So this piece takes the opposite position to most of that material: a feature store solves three specific problems, and if you do not have at least one of them today, you do not need one yet. Standing it up early buys you a dependency, an operating cost, and something new to own, in exchange for solving a problem you did not have.
What a feature store actually does, and what it does not
Strip away the positioning and a feature store does one thing. It computes a feature value once, stores it, and serves it to anything that needs it, from a single shared definition.
That is less obvious than it sounds, because features get computed in at least two places in every ML system. The training pipeline computes them over historical data to assemble a dataset. The serving path computes them over current data to score a live request. In most organisations those are two codebases, written by two people, at two different times, against two different data sources. A feature store collapses that into one definition with two read paths.
What it does not do matters just as much, because most of the overbuying happens here.
- It is not a data warehouse. The warehouse holds raw and modelled data. The feature store holds the derived values a model actually consumes, which is a much smaller and more opinionated set.
- It is not a cache. A cache stores whatever you put in it and has no notion of correctness over time. A feature store knows what a feature means, when it was computed, and which version of it a given model trained against.
- It does not make your model better. It stops your model degrading for reasons that have nothing to do with the model. That is a real benefit and a much less exciting one.
- It is not a prerequisite for production ML. A great deal of working production ML runs without one, correctly.
That last point is where most of the confusion sits. A feature store is not a maturity badge. It is a response to specific symptoms, and the useful question is whether you have the symptoms.
The three problems that justify one
Three problems. The honest test is whether you recognise one of them in your system right now, not whether you might encounter it at some future scale you have not reached.
1. Training-serving skew. The model trained on features computed one way, usually in a batch job over the warehouse. Production serves those same features computed a different way, usually in application code against a live database. The definitions drift, often by something as small as how nulls are handled or whether a window is inclusive. The model is then scoring inputs it never saw during training, and accuracy degrades for reasons no amount of retraining will fix, because the model is not what is broken. This is the single most common cause of a model that validated beautifully and underperformed in production, and it is the problem a feature store was invented to solve.
2. Duplicated feature logic across models. Your fraud model computes a thirty-day transaction average. So does your churn model. So does the credit risk model, each in its own codebase, each subtly differently. When the business changes what counts as a transaction, someone has to find all three and fix them consistently. They will not. The cost here is not compute, it is the slow accumulation of definitions nobody can reconcile, and the growing reluctance to change anything because the blast radius is unknown.
3. Online serving latency. The feature exists, but getting it takes too long. Recomputing a rolling aggregate inside the request means querying transactional tables under load, and that will not fit a tight budget. As covered in our piece on real-time AI architecture, the feature fetch typically gets 10 to 15 ms of a 100 ms interactive budget, and recomputation routinely costs an order of magnitude more than that. An online store turns an expensive computation into a single key lookup.
One of these is enough. Two makes it urgent. If you have none of them, read on to the section on deferring, because there is a cheaper intermediate step most teams skip past.
Online and offline: the two halves people conflate
A feature store is really two systems wearing one name, and knowing which half you need is what separates a sensible purchase from an expensive one.
| Half | Serves | Optimised for | Typically runs on |
|---|---|---|---|
| Offline store | Training and backtesting | Large scans, point-in-time correct joins | Your existing warehouse or lakehouse |
| Online store | Live inference | Single-digit millisecond key lookups | Redis, DynamoDB, or similar |
The offline half carries the capability teams consistently underestimate: point-in-time correctness. When you build a training dataset, each row has to reflect only what was knowable at that historical moment. Join a customer’s current lifetime value onto a transaction from eight months ago and you have leaked the future into your training data. The model learns from information it will never have at inference, validates far better than it should, and disappoints in production. Doing this correctly with ordinary SQL joins across many features is genuinely difficult, and it is the part of a feature store that is hardest to replicate yourself.
The online half is conceptually simpler. It holds the current value per entity and answers lookups fast. The engineering is in keeping it fresh and consistent with the offline half, which is where scalable data pipeline architecture stops being a background concern and becomes the thing the whole system rests on.
Here is the practical consequence. Plenty of teams have a skew or duplication problem and no latency problem at all, because they score in batch. Those teams need the offline half and nothing else. Buying a platform that does both, and paying to operate an always-on online store nothing reads from, is the most common way to overspend on this.
When you do not need one yet
Five conditions. If most of them describe your situation, a feature store is solving a problem you have not got.
- You have one or two models in production, not a portfolio.
- Each feature is used by exactly one model, so nothing is being duplicated.
- You score in batch, so no feature has to be read in milliseconds.
- The same team writes both the training and the serving code, so definitions drift slowly if at all.
- Your features are mostly direct column selections rather than complex time-windowed aggregations.
Teams in that position who buy a feature store get a new system to operate, a migration that consumes a quarter, and no measurable improvement, because the problems the tool addresses were not present.
There is a middle option that gets skipped far too often. Extract your feature definitions into a shared library, and materialise them as scheduled tables in the warehouse. Training reads the tables. Serving imports the same library. That removes duplication and most of the skew risk, costs close to nothing, adds no new infrastructure, and takes days rather than a quarter. It will not give you millisecond reads, and it will not give you proper point-in-time correctness. But it buys you a year of clarity about whether you genuinely need those things, and that is usually the decision you are actually trying to make.
Build, buy, or defer: a decision framework
| Option | Fits when | Real cost | Main risk |
|---|---|---|---|
| Defer | Duplication only, batch scoring, small portfolio | Days of engineering, no new infrastructure | No point-in-time correctness, so skew can still creep in |
| Buy managed | Need the online half, several models, no platform team | Licence plus integration, weeks to first value | Pricing scales with feature count, and lock-in is real |
| Build on open source | Platform team exists, warehouse and key-value store already run | Lower licence spend, materially higher operating load | You now own availability and upgrades for a critical path |
| Build custom | Requirements genuinely unmet by existing options | Quarters, not weeks | Rebuilding point-in-time correctness badly |
A note on the last row. Teams who build custom almost always underestimate point-in-time correctness and online and offline consistency, and those are precisely the two things existing tools have already solved properly. If your reason for building is that the managed options are expensive, build on open source instead. If your reason is that your requirements are unique, test that assumption hard before committing quarters to it, because it is usually the data platform underneath that is unusual, not the feature serving. That distinction is worth resolving before you commit, and it is the kind of question our data platform engineering work starts with.
What it actually costs to run
The licence is the part everyone models, and it is rarely the part that hurts.
- Online store infrastructure. Always-on, sized for peak, and it costs the same at 3am as at noon. This is a standing bill from the day you turn it on.
- Streaming or scheduled pipelines. Keeping the online store fresh runs continuously whether or not anyone is scoring anything. Freshness is a recurring cost, not a one-off build.
- Integration into existing pipelines. Every model you migrate is real work, and the second one is not much cheaper than the first.
- Ownership. The one that gets missed.
That last line is the one worth dwelling on. A feature store is an internal platform product, which means it has users. Someone has to curate definitions, onboard teams, answer questions, handle the pager, and decide what gets deprecated. Without a named owner it degrades into exactly the inconsistent sprawl it was bought to prevent, except now there is also a bill. Standing it up and operating it properly is MLOps and AI infrastructure work, and treating it as a one-off project rather than an owned platform is the most reliable way to waste the investment.
Budget a person, not just a line item. If you cannot name the owner, you are not ready to buy.
Also read
- Real-time AI architecture: what actually works at enterprise scale
- From pilot to production: the MLOps lifecycle behind scaling AI
- Why scalable data pipelines are the backbone of modern enterprises
This article is about the decision. For the governance side, meaning how a feature store underpins discovery, reuse, and audit across an ML platform, that sits within the broader pilot-to-production guide.
The bottom line
A feature store is not a milestone on a maturity curve. It is a response to three symptoms: training and serving disagreeing about what a feature means, the same logic reimplemented across models, and features that cannot be computed fast enough at inference. Have one of those and it earns its place quickly. Have none and the cheaper move is a shared definition library and scheduled tables, which buys you a year of evidence before you commit to infrastructure.
Decide on the symptoms you have, not the architecture diagram you would like to show. And before you sign anything, name the person who will own it.
Work out whether you need a feature store
Bring us your model portfolio and your serving path. We will tell you whether skew, duplication, or latency is the real problem, and what the cheapest fix looks like.