Ask Claude, Gemini, or ChatGPT what size you wear in a specific brand's jacket. You will get an answer. It will be fluent and confident, and it will often be wrong.
That gap between fluency and accuracy defines the frontier of AI in apparel sizing. These models reason well over language. Sizing is a measurement problem described in language, and a language model only has the description.
The stakes are measurable. Coresight Research puts the average US online apparel return rate at 23.4% in 2025. Nearly 70% of shoppers who returned clothing online cited size and fit as the reason. AI agents are taking over more of the discovery and purchase journey, and the question is shifting from whether models will answer sizing questions to what data backs those answers.
How these models fail here is well documented and structural, not a one-off glitch. OpenAI's own research found that language models hallucinate because training and evaluation reward guessing over acknowledging uncertainty. Pretraining maximizes the likelihood of the training data, which rewards plausibility over accuracy, and when the underlying data is thin, plausible falsehoods follow naturally. Post-training rarely removes the behavior, because most benchmarks score a wrong answer and "I don't know" the same way.
Apply that to sizing. A model asked for your size in a specific jacket sits in exactly that low-data condition, and it's built to produce an answer rather than decline. The output carries the same confident tone whether it's right or not.
A shopper who gets no recommendation hesitates. Coresight found that roughly 40% of US shoppers have abandoned an online apparel purchase due to confusing or missing product information. A shopper who gets a wrong, confident recommendation often does something costlier. They order two sizes and return whichever one doesn't fit. Fluency without ground truth yields higher returns.
A frontier model learns from publicly available text and images. That collection contains size charts, which are marketing artifacts rather than engineering documents, with inconsistent units and unlabeled measurement points and no reliable way to tell whether a chart describes a body or a garment. Decades of vanity sizing sit in that collection too. A size 8 twenty years ago and a size 8 today describe different bodies. Forum threads sit there as well, with one shopper saying a shirt runs small and another saying it runs large. Both shoppers can be right. Bodies vary that much from person to person.
What the collection does not contain is human body measurement data at a meaningful scale. The most comprehensive public dataset on body size and shape is ANSUR II, the US Army's 2012 anthropometric survey, which includes 93 directly measured dimensions from 6,068 soldiers and was released in 2017. It's excellent data, but it does not represent a consumer population. The Army itself notes the sample represents the US Army at the time of collection and may not represent other populations of interest. Its commercial civilian counterpart, CAESAR, covers roughly 4,400 North American and European adults and carries a license fee.
Combine those two datasets, and you get roughly 10,000 people in the public and commercial anthropometric record. Most were soldiers, and the data is over a decade old. That's the body data foundation available to any general-purpose model.
There's a second gap worth noting. Even ANSUR II withholds its 3D whole-body scans from public release to protect participant privacy. The most detailed body data in the most open dataset in this field stays unpublished. That's how this domain usually works.
Bold Metrics has built a proprietary body data library since 2014. That library now holds 250M+ digital twins and 12B+ body data points, with 750M+ fit simulations run against them. Each twin carries 50+ body measurements, determined from four to six simple shopper inputs.
Set that against the roughly 10,000 people in the public record. The gap runs close to five orders of magnitude. That's why a purpose-built platform and a general-purpose model produce different answers to the same question.
Building that library took ten years. Bold Metrics trained its AI models on one of the largest cumulative datasets of high-fidelity human body scans in the world, paired with proprietary machine learning algorithms proven across thousands of fittings. Combined with garment data, purchase and return behavior, and customer body data, the system pinpoints body measurements with remarkable accuracy, generating a digital twin. The models keep improving, continuously refined using real-world purchase and return data.
Learning how bodies vary. Determining 50+ measurements from a handful of inputs requires knowing how bodies actually vary, including how they vary within a single height and weight. Two shoppers at 5'10" and 180 pounds can need different sizes in the same shirt. That variance lives in the distribution, and a distribution only comes from a very large number of real bodies.
Coverage at the edges. Averages are easy to model. The commercial value sits at the tails, with shoppers who fall between sizes or sit outside the range a brand originally graded for. Those shoppers drive a disproportionate share of returns, and they convert at a lower rate. A decade of data puts density where it counts.
Validation against outcomes. Every recommendation the platform makes can be checked against what actually happened next, from purchase through return. That feedback loop separates a model that sounds right from one that's been shown to be right, and it's why accuracy keeps improving instead of leveling off.
Body data alone doesn't produce a size recommendation. The garment side matters just as much, and it's even less public.
Coresight's May 2026 framework makes a similar point from the brand side. The quality of sizing recommendations reflects the sizing intelligence behind them, and structured garment measurements and grading rules need to be anchored in consistent standards before any recommendation engine can work, even for new styles with no sales history.
That information lives in tech packs, on internal servers, under NDA. It differs by brand and season, and often by individual style within a line. Bold Metrics ingests it and maps it against each shopper's digital twin. Virtual Sizer™ handles the mapping from body measurements to garment specs. Smart Size Chart™ delivers that mapping on the storefront, showing how each size fits a specific shopper instead of returning a bare label. Apparel Insights® sends the aggregate body data back to design and merchandising teams. Future size runs get graded against the customers a brand actually has.
No general-purpose model has access to any of this. That data belongs to brands and remains subject to agreements that exist for good reason.
Body measurements are personal, and brands need to treat them that way. Shoppers share their measurements with a particular brand for a particular reason. That creates real obligations, including consent-based collection and a documented chain of custody that covers how data is stored and retained.
Under GDPR, measurements tied to an identifiable person count as personal data and carry the full set of lawful basis and data minimization duties. Body scans push further into territory that biometric statutes were written to cover.
A shopper who pastes measurements into a general-purpose chat assistant moves that data outside a brand's control. A brand routing fit questions through an unvetted third-party model carries a risk it likely hasn't priced. A purpose-built platform can be designed around consent and data handling from the start. A model trained on the open web can't acquire that posture after the fact.
Agentic commerce is arriving fast. A December 2025 Coresight survey found 58% of US consumers familiar with AI had used or intended to use AI tools for shopping. In that world, the model becomes the interface, and the interface still needs a source of truth to call on.
Coresight warns that brands with inconsistent or incomplete sizing data receive lower recommendation confidence and see errors repeated at scale, which, over time, can lead to being deprioritized as high return rates generate negative performance signals. Wrong recommendations delivered at machine speed don't stay a data problem for long. They turn into a visibility problem.
Bold Metrics built Agentic Sizing Protocol™ for that handoff. ASP lets an agent query real body data and real garment specs and return an answer grounded in both. Retailer credentials stay out of the agent's reach and never appear in a model response. The model runs the conversation while the platform supplies the measurement.
These figures come from live client deployments rather than projections:
They trace back to a body data library built over ten years, paired with brand-specific garment data and a feedback loop that keeps tightening, and no prompt produces that.
Frontier models will reshape how apparel gets discovered and purchased, and they're worth building for. They aren't a substitute for proprietary data built up over ten years. Language models return plausible answers. Fit intelligence needs accurate data, and accuracy comes from real bodies and real garments, validated against outcomes.
Ready to see what a decade of digital twins does for your return rate? Explore the Bold Metrics platform or request a demo.