OutfitScore Research · Report No. 05 · September 2026

The Neglected Hem

What 20,814 AI-scored outfits reveal about pooling, break, and unfinished trousers

Abstract I classified free-text improvement tips on 20,814 scored fashion and complete analyses processed between 27 September 2025 and 24 September 2026. Mean overall score is 66.5 (median 66.0). Keyword flags on areas_for_improvement show that hem pooling / dirty break is the most common unpublished finish complaint: 1,016 rows (4.9%). Wrinkle or press language appears on 371 rows (1.8%); visible-sock language on 196 (0.9%). Accessory language (21.3%) and texture language (2.9%) already have dedicated OutfitScore posts and are used here only as contrast. Hem-flag rate peaks in the 60–74 band (621 / 9,538 = 6.5%). This is an observational classification of model tips, not a treatment study: hem-flagged rows mean 67.7 against a corpus mean of 66.5. No causal score lift is claimed.
Blue jeans meeting black sneakers at the ankle on pavement, the hem-to-shoe line this report measures
Figure 1. The meeting this report counts: denim hem against the shoe. Stock photograph, not a user upload.

Fashion media will sell you a necklace, a texture story, or a French tuck. Those are real levers. OutfitScore already wrote them up, because they dominate the tip stream. This report is about the lever that is common, cheap, and still unpublished: the place where the trouser meets the shoe.

Pooling is not a taste argument. It is a geometry argument. Extra fabric at the ankle breaks the vertical line, hides the shoe, and reads as unaltered even when the rest of the outfit is considered. The model says so in plain language — “hem the jeans to a proper break,” “eliminate pooling,” “clean line from knee to shoe” — on 1,016 scored rows. That is enough volume to publish. It is not enough volume, or the right design, to pretend that flagging a hem causes a higher score.

20,814Scored outfits
1,016Hem flags
4.9%Of corpus
6.5%Rate in 60–74
How to read the numbers Every headline figure in this report is frozen on 24 September 2026. The page does not live-query production. If the corpus grows tomorrow, these n= values stay. Cite the snapshot, not “current site stats.”

Section 1Methodology

The unit of analysis is one row in the OutfitScore analyses table. I keep a row if analysis_type is fashion or complete, analysis_result is present, and a numeric overall score can be read from overall_score, score, or total_score. Roast, makeup, body-type, color-season, and accessory-only analyses are excluded. Soft-deleted rows are not filtered: a deleted analysis is still a real observation of how the model talked about an outfit.

The window is not a rolling 30-day panel. It is every qualifying row then in production: 27 September 2025 through 24 September 2026. That is 363 days, not a season. Seasonal mix (shorts vs wool, boots vs sneakers) is therefore averaged, not isolated.

Score extraction follows the same numeric-regex guard used by seo_stats_service. Non-numeric strings such as “N/A” become null and drop the row from the scored corpus. After that filter, n = 20,814. Mean 66.5, median 66.0. A potential score is present on 9,319 of those rows; the mean gap between potential and current is 15.0 points. That gap is a model-implied ceiling, not a measured before/after.

How a “flag” is defined

OutfitScore does not store a structured enum for “hem pooling.” The model writes free text into areas_for_improvement. I classify that text with case-insensitive regular expressions. quick_wins is ignored so a locked marketing field cannot inflate counts. A row may match more than one flag. Presence is not exclusive unless a table says so.

The hem pattern requires pooling, puddling, ankle bunching, “clean break” / “proper break,” or an explicit instruction to hem jeans, trousers, pants, or chinos. It is deliberately tighter than a naive search for the word “hem,” which would also catch skirt hemlines and “hem the t-shirt.”

pool(ing)? | puddl(e|ing) | bunching at the ankle
clean break | proper break
hem(med|ming)? the (jeans|trousers|pants|chinos)
have the (jeans|trousers|pants) … hem

Wrinkle language requires wrinkle, unpressed, iron, or steam-to-remove. Sock language requires no-show socks, visible socks, sock length, or white socks. Accessory language (necklace, chain, watch, bracelet, jewelry, “accessor”) and texture language (“textural variety,” “lack of texture,” “add texture”) are contrast flags only.

Publication floor A rate is printed only if the flag n is at least 100. Visible tags (33), lint or pilling (78), scuffed shoes (28), and uneven sleeve rolls (4) exist in the corpus and are named as too thin to rate. Dimension sub-scores on hem-flagged rows (fit / styling / details) have n = 69 and are likewise unpublished.

False positives are real. “Cuff” can mean a sleeve or a jean roll. “Break” can mean a suit break or a photographic pause. The regex tries to keep those out; it will not catch every miss, and it will not refuse every overlap. The method is keyword classification of model prose, not a human-coded content analysis with inter-rater kappa. Report 01 used the same family of method on a 492-row window. This report is larger and narrower: one unpublished finish, not five deficiency clusters.

Section 2What this report is not

Three OutfitScore posts already occupy the high-volume “neglected detail” shelf. Cloning them would be a second URL, not a second finding.

Already published This corpus (n=20,814) This report
The Accessory Gap 4,434 (21.3%) Contrast only
Texture Variety 598 (2.9%) Contrast only
The French Tuck 240 (1.2%) Contrast only
Hem pooling / dirty break 1,016 (4.9%) Headline

Accessories remain the loudest single family of tips. That is consistent with Report 01, which found accessory absence in 20.5% of recommendations on a much smaller 2026 spring window. Texture is smaller here than in that 492-row study because the regex is tighter (“textural variety” rather than every mention of knit). French tuck is already a named post and a named move. None of those three is “neglected” in the editorial sense. The hem is.

This is also not a how-to for bleaching denim, dyeing fibre, or religious dress. Modest coverage and hijab geometry belong to a different guide. The hem claim here is secular and mechanical: extra fabric on the shoe, or not.

Section 3The unpublished finish stack

Once accessories, texture, and the named tuck are set aside, three finish flags clear the n ≥ 100 floor.

Finish flag n % of 20,814 Mean score (observational)
Hem pooling / dirty break 1,016 4.9 67.7
Wrinkle / press / steam 371 1.8 71.3
Visible / no-show socks 196 0.9 66.7
Corpus 20,814 100 66.5

Hem pooling is more than twice wrinkles and five times visible socks. That ranking is the finding. Mean scores on flagged rows sit near or above the corpus mean. That is the opposite of a naive “this mistake tanks you” story, and it is why a causal +N claim is refused.

Two mechanisms explain the inversion without requiring a treatment effect. First, the model’s improvement list is capped by score band: high-scoring outfits get fewer tips, and those tips get pickier. A 72 that is otherwise assembled will be told to hem. A 48 will be told to change the garment, the occasion, or the silhouette. Second, pooling is easier to see on a full-length photograph of trousers that almost work. Cropped selfies and outfit-from-the-waist-up shots under-count hems; they also tend to score differently. Photo type is a separate study.

The model does not deduct seventeen points for a puddle. It spends a scarce tip slot on the puddle when the rest of the outfit is already in the conversation.

A union of hem, wrinkle, socks, tag, lint, scuff, and uneven sleeve is present on 1,687 rows (8.1% of the corpus; 9.5% of the 60–74 band). That union is a sensitivity check, not the headline. The hem alone is the unpublished volume story.

Section 4Where pooling shows up

Band sizes in the scored corpus: below 60, n = 4,852; 60–74, n = 9,538; 75–84, n = 3,711; 85 and above, n = 2,713. Hem-flag counts that clear n ≥ 100:

Hem-flag rate by score band Published bands only where flag n ≥ 100. Dashed line = corpus rate 4.9%. 8% 6% 4% 2% 0 4.9% 3.1% 151 flags <60 n = 4,852 6.5% 621 flags 60–74 n = 9,538 4.2% 155 flags 75–84 n = 3,711 unpublished flag n = 89 85+ n = 2,713 Mid-band peak: the model still has tip budget and the outfit is close enough to finish. 85+ rate withheld. Hem-flag rate (n ≥ 100) Corpus 4.9% 85+ unpublished
Figure 2. Hem-flag rate is highest in the 60–74 band, where the model still has tip budget and the outfit is close enough to finish. The 85+ rate is unpublished (flag n = 89).

Wrinkle flags, by contrast, are more common in 75–84 (119 / 3,711 = 3.2%) than in 60–74 (141 / 9,538 = 1.5%). That is the picky-tip pattern in a different costume: once fit and accessories are not the story, the model starts talking about press. Sock flags stay near 1% in every published band and never dominate.

The mid-band concentration is consistent with the engine’s own cap rules, documented in the analysis prompt: scores 85+ are limited to one improvement tip; 65–84 get two; below 65 may get three, ranked by expected impact. Below 60 those three slots are expensive. The model spends them on fit, occasion, and cohesion. Pooling loses the auction. In the 60s, pooling wins more often because the outfit is already “about” trousers and shoes.

Section 5What the tips actually say

The language is repetitive in a useful way. It is not “buy new jeans.” It is “change the length of the jeans you already photographed.” Representative phrasings, truncated, from hem-flagged rows:

  • Hem the jeans to a proper break to avoid bunching at the ankle and elongate the leg.
  • Have the pants professionally tailored to a clean break; improves the vertical line and polish.
  • Tailor the denim to eliminate pooling, creating a cleaner line from knee to shoe.
  • Have the cargo pants hemmed to a slight break to eliminate bunching at the ankle.
  • Hem the jeans or cuff them to eliminate pooling.

Two interventions sit inside that language. One is a tailor: a permanent inseam for the shoe you actually wear. The other is a cuff or a crop, which is reversible and often good enough for sneakers. The model treats both as hem work. This report does not split them. A cuff that still stacks is still a puddle.

The shoe is not optional in the geometry. A jean hemmed for a runner will break wrong on a boot. That sentence already appears in OutfitScore’s jeans-for-body-type guide as styling advice. Here it is an empirical note: the flag is about the meeting of hem and shoe, not about the jean in isolation.

Practical reading, not a lift If you are in the 60s and the photograph is full-length, look at the ankle before you look at a new top. The corpus cannot tell you that hemming will add points. It can tell you that this is the finish the model still bothers to name after accessories have had their say.

Section 6Why “neglected” is the right word

Neglected, here, does not mean rare. It means editorially skipped. Search and social already know the accessory gap and the French tuck. They do not know, as a citable n=, that nearly one in twenty scored OutfitScore outfits is told to deal with fabric on the shoe.

It is also neglected in the dressing sequence. People buy the jean for the rise and the wash. They try it on with the shoes they wore to the shop, or in socks on a carpet. They wear it with a different shoe. The puddle is a leftover of that sequence, not a style choice. Wrinkles are a morning sequence leftover. Visible athletic socks under shorts are a default leftover. Finish flags are leftovers. Fit flags are usually purchase errors. That is why a hem is cheaper than a new silhouette even though this report will not price it in points.

Report 01, on 492 outfits, already listed “trouser hems that pool” as a Tier-2 mechanically correctable issue that requires no new clothing. This report is the n= behind that sentence, a year later, on a corpus forty times larger, with a published regex and a refusal to convert the flag into a fake lift.

Section 7Cuff versus tailor

The tips collapse two different jobs into one flag. A tailor shortens the inseam to a shoe. A cuff rolls the extra length so the shoe can be seen tonight. Both can produce a clean break. Both can fail. A thick double cuff next to a blazer still reads as an afterthought; OutfitScore’s menswear guide already says so. A permanent hem cut for a 40 mm runner will stack on a Chelsea boot. The flag does not know which shoe will be worn tomorrow. It only knows what the photograph showed.

That is why this report will not rank “cuff” against “tailor” as interventions. The regex cannot see a receipt. What it can see is that the model’s preferred verbs are hem, tailor, and eliminate pooling. Cuff appears as a fallback, often in the same sentence as a professional hem. If you only do one reversible thing before a photograph, a single clean cuff that shows the vamp is the version of the finding you can test without a needle. If the jean is a keeper, the needle is the version that survives a different shoe.

The 17-point “potential” article on this site names footwear, a French tuck, and one focal accessory as the stacked close of a typical gap. This corpus’s mean potential gap is 15.0, on a later and larger window. Hem pooling is not in that trio because it is not the modal tip. It is the modal unpublished finish tip. Readers who already added the chain and the tuck and still sit in the 60s are the population Figure 2 is about. The leftover at that point is often fabric on the shoe, not another necklace.

Section 8Limitations

The corpus is voluntary. People who upload an outfit to an AI rater are not a random sample of dressers. They are more likely to photograph a full look, more likely to be curious about a number, and more likely to wear the kind of trouser that can pool — jeans, chinos, wide-leg. Skirts, shorts, and cropped trousers are under-counted in the hem flag by construction.

The classifier is a regex on model English. Multilingual tips, metaphors, and novel phrasings are missed. Overlap with cuff-as-style and break-as-suit-term is possible. I did not double-code a human gold set for this report. A different taxonomy would move rows between hem, fit, and footwear.

Photo crop is unmeasured here. A mirror selfie that stops at the knee cannot draw a hem flag. If mid-band looks are more often full-length, part of the 6.5% rate is a photograph effect. Report-style photo-type work exists elsewhere on the site; it is not merged into these denominators.

Potential-score gap on hem-flagged rows is 14.3 against 15.0 on the corpus among rows that have a potential field. That is not a lift. It is almost the same gap. Publishing it as “hemming closes the gap” would be a misuse of the field.

Thin flags are named and withheld. Absence of a published rate is not evidence that tags, lint, or scuffed shoes do not matter. It is evidence that this snapshot cannot support a percentage without looking precise.

Section 9Conclusion

On 20,814 scored fashion and complete analyses, the most common unpublished finish flag is hem pooling and a dirty break: 1,016 rows, 4.9% of the corpus, 6.5% of the 60–74 band. Wrinkles and visible socks trail it. Accessories and texture still dwarf it and already have posts. The honest sentence for a journalist is not “hemming raises your score by N.” It is: when the model still has something small left to say about a mid-band outfit, the thing it names at the shoe is extra fabric.

If you want the engine to look at your own hem, the tool is the same as it was: the free AI outfit rater. This page is the study. Pitch this URL.

How To Cite This Report

Title: The Neglected Hem: What 20,814 AI-Scored Outfits Reveal About Pooling, Break, and Unfinished Trousers
Author: Saad, Founder of OutfitScore
Publication: OutfitScore Research Reports, No. 05
Date: September 24, 2026
Sample: n = 20,814 scored fashion/complete analyses, 27 Sep 2025 – 24 Sep 2026
Saad. (2026). The Neglected Hem: What 20,814 AI-Scored Outfits Reveal About Pooling, Break, and Unfinished Trousers. OutfitScore Research Reports, No. 05. Retrieved from https://outfitscore.com/research/the-neglected-hem