Predictive analytics in real estate: does the seller data actually work?
Predictive analytics ranks homeowners by how likely they are to sell, but only about 5% of homes sell in a given year, so even the best-scored lists are mostly non-sellers. Vendor accuracy claims are unaudited, the scores cost extra on top of postage, and the results come from mailing a ranked list consistently, not from the algorithm alone.
What predictive analytics actually does
Predictive analytics scores homeowners on how likely they are to sell. HousingWire describes it as a method that reads historical data, current trends and machine learning to "forecast future outcomes and consumer behaviors." In plain terms, a vendor runs a model over the properties in your market and hands you a ranked list, with the owners it thinks are closest to listing near the top.
The inputs are the kind of signals that historically line up with a sale: how long someone has owned the home, their age and life stage, how much equity has built up, the loan situation, how much the home has appreciated. SmartZip says its model weighs more than 25 data points across 122 million properties. Top Producer markets its version as flagging the top 20% of households most likely to sell within 12 months. The promise is the same across all of them: mail fewer people, close more deals.
This is general information, not financial advice. Any return you get from a lead source depends on your market, your data quality, your offer and your follow-through, so treat vendor claims as marketing and consult a professional before you commit real budget. The figures below are ranges and vendor claims, not promises.
The base rate that decides everything
Before any model earns a dollar, one number sets the ceiling on what it can do: how often homes actually sell. In 2024 there were 4.06 million existing-home sales in the United States, the lowest annual total since 1995, per National Association of Realtors data reported by the NAHB. The Census Bureau counts about 148 million housing units, roughly two-thirds owner-occupied. Run the division. Somewhere near 1 in 20 owned homes changes hands in a year, and that is the ceiling every model works under.
SmartZip, a vendor that sells predictive leads, puts the national turnover baseline at 5% in its own marketing, which lines up. So the honest starting point is this: pick any homeowner at random and there is about a 95% chance they will not sell this year. A model's whole job is to beat that 5%. By how much is the only question that matters, and whether the lift is worth what the scores cost on top of your postage.
What the vendors claim
The headline numbers come from the companies selling the scores, so read them as marketing until someone independent checks them. SmartZip claims it "predicted 72% of home listings" in a given year, and that agents mailing only its highest-scoring prospects saw turnover of 27% against that 5% national baseline, which it frames as 4.6 times better. It also notes that only about 4% of properties reach its top score tier.
Take the best of those claims at face value and the math still bites. A 27% turnover on the top tier means 73 out of every 100 owners you pay for and mail this year will not sell this year. That is the good case, on the top 4% of the market. HousingWire, reviewing the category, is blunt about the weak spot: "The data is only as reliable as the sources used." No vendor publishes an independent, audited hit rate, and you should notice which number is always missing from the sales page.
Tools and pricing
The pricing below is from HousingWire's 2026 roundup. Treat it as a starting point; most of these quote by territory or ZIP and negotiate.
| Tool | Roughly | Built for |
|---|---|---|
| SmartZip | ~$500/month | Agents farming a territory for listings |
| Top Producer | $179/month and up | Agents wanting scores inside a CRM |
| Fello | $415/month and up | Database reactivation and seller scoring |
| PropStream | $99/month and up | Investors pulling lists and property data |
| Revaluate | Custom | Move-likelihood scoring on a database |
SmartZip at roughly $500 a month is $6,000 a year for a farm. For a listing agent who closes even two extra deals from it, that pays for itself several times over. For a wholesaler working on thin per-deal margins and heavy mail volume, paying a premium for scores on top of postage is a different calculation, and often a losing one.
The math on a predictive list
Here is where the promise meets the postage bill. Say you buy a 500-home predictive list and it converts at 8% over a year, which is above the 5% base rate and below SmartZip's best-case 27%. That is 40 sellers hiding in the list. You would love to mail only those 40. You cannot, because the score tells you a probability, not a name and a date. So you mail all 500, several times, to reach the 40.
At around a dollar an all-in mailed piece and six touches over the year, that 500-home list costs roughly $3,000 in postage and print, on top of whatever you paid for the scores. Land 40 sellers and even a handful of closings makes it work. Convert at 4% instead, barely above random, and the same spend chases 20. The scores did nothing. The model's value is not the scores themselves. It is whether the lift over the base rate is big enough to cover the extra cost of buying it.
The mistake almost everyone makes
The common advice is to buy the highest scores and mail only those. For most people that is wrong, and the reason is in SmartZip's own data. Only about 4% of properties hit the top tier, and even those sell at 27% in the best case. A home scored 900 and a home scored 600 both mostly do not sell this year. Both are long shots. The difference is that the 600s are a far larger pool, and a real chunk of them will still list.
Chase only the top 4% and you are mailing a tiny group very hard while ignoring a much bigger group that produces real listings at a lower rate. The better move is to rank a well-defined market and mail the top slice consistently over 6 to 12 months, widening the slice to what your budget covers rather than shrinking it to only the highest scores. Timing beats intensity. A model tells you who is warmer, not who is selling next Tuesday, so you have to be in the mailbox when the life event that triggers a sale finally lands.
Why timing is the part that trips people up
A model that says an owner is likely to sell within 12 months is not telling you which month. The sale gets triggered by something the data cannot schedule: a job transfer, a new baby, a divorce, an aging parent, an equity number crossing a line in the owner's head. You cannot predict the week. You can only make sure your name is the one in front of them when the decision finally lands.
That is why cadence beats precision. An owner who scores high in January might not list until October, and a single postcard mailed in February is long forgotten by then. Mail the same ranked list every 6 to 8 weeks and you stay present across the whole window instead of betting everything on one send. The people who win with predictive data are rarely the ones with the best scores. They are the ones who mailed the list a dozen times while the competition quit after two.
Predictive scores vs trigger lists
There are two ways to guess who will sell, and they are not the same product. A predictive score is statistical. It reads dozens of quiet signals and estimates a probability. A trigger list is behavioral. It catches an event that already happened and usually pushes an owner toward selling. Probate filings, pre-foreclosure notices, tax delinquency, code violations, expired listings and out-of-state absentee owners are trigger lists. The event itself is the signal, and there is nothing to model.
Trigger lists convert harder because the motivation is real and present, but they are small and every investor in the county is mailing them. Predictive scores cover the whole market and can flag owners before an obvious trigger appears, which is their genuine edge, but the motivation is softer and the false-positive rate runs high. Most serious operators run both: trigger lists for near-term deals, and a ranked predictive or equity list for volume and for reaching sellers earlier than the competition. One does not replace the other, and a vendor that tells you its score makes trigger lists obsolete is overselling.
Where predictive earns its keep, and where it does not
It fits listing agents best. If you farm a neighborhood and your cost per closed listing is thousands of dollars in commission, a $6,000 model that surfaces a few extra sellers a year is easy math. It also pairs well with a farming plan you were going to run anyway.
It fits wholesalers and investors less cleanly. Your margins are thinner, your mail volume is higher, and a generic seller score is not tuned to distress or motivation. A ranked list of likely sellers is a fine starting universe, but you still filter it against the signals that matter to you: equity, absentee status, condition. For many investors a well-built motivated-seller list and disciplined mail beats a pricey seller score they cannot fully use.
What actually moves the needle
Scoring is half the job, and the cheaper half to talk about. The other half is putting a piece of mail in front of the ranked owners, again and again, until one of them is ready. A perfect list you mail once does nothing. A decent list you mail six times over a year produces deals. The value shows up in the follow-through, not the algorithm.
This is the part Farmrix is built around. It scores every owner in your market on how likely they are to sell in the next 6 to 12 months, ranks them, then prints and mails the postcards to the top of that list, so the scoring and the mailing are one job instead of two vendors and a spreadsheet. Packages start at $1,195. If you already own a scoring tool and a reliable print-and-mail pipeline, keep using them. If you are paying for scores and then wrestling with mail separately, closing that gap is usually where the wasted spend is.
How to buy it without getting burned
Ask every vendor for the number they never lead with: what share of the people you mail actually list within 12 months, measured, not modeled. Push on the base rate, since anything near 5% is just the market. Start with the smallest territory they will sell so you can measure your own conversion before you scale. Commit to at least six touches over a year before you judge it, because one mailing to any list is a test of nothing. Then compare the cost per closed deal against a plain ranked mailing list run with the same discipline, whether that is a standalone score plus your own mail or a combined tool like Farmrix that scores, ranks and mails in one pass. If the scores do not beat that, you were paying for a story, not a lift. Track it by cost per closed deal over a full year, not by open rates or how good the list looked in a demo.
Get the next guide
One practical email when we publish. No drip sequence, no pitch.