OutfitScore Research · Report No. 07 · October 2026

The Context Gap

What 20,883 outfit scores reveal about worn looks vs individual pieces

Abstract I scored 20,883 fashion and complete analyses processed between 27 September 2025 and 2 October 2026. Mean overall score is 66.3 (median 66.0, SD 15.1). On exact item_type strings that clear n ≥ 100, complete_outfit means 68.6 (n = 13,469) against individual_pieces 57.4 (n = 957) and flat_lay 64.1 (n = 218). The worn-versus-pieces gap is 11.2 points. The below-60 rate is the sharper cut: 45.4% of exact pieces versus 19.2% of worn looks. Pieces are not a tight 55–72 band: only 42.9% land there, and the piece SD is 19.6 against 14.5 for worn looks. Labels are exact string matches, not merged families. This is an observational snapshot of model labels, not a treatment study. No causal score lift is claimed.
Woman in a red coat walking on a city street — a worn full look, the context this report measures
Figure 1. The unit this report measures: a worn full look, not a hanger shot. Photograph by Tamara Bellis on Unsplash (Unsplash License). Stock photograph, not a user upload.

OutfitScore already published a how-to on individual pieces versus full outfits. That post used 86 piece analyses and 707 complete looks, and it claimed pieces cluster in a 55–72 band. This report is the n= behind that shelf: what the model actually writes when the label is complete_outfit, individual_pieces, or flat_lay, on a corpus more than twenty times larger.

The tempting headline is that a lone garment cannot score high. It can. Exact pieces reach 95. The honest headline is that they also reach 0, that nearly half of them sit below 60, and that a worn look and a piece do not share a band. The April blog's tight mid-band is false on this snapshot. The how-to stays a how-to. This page is the study.

20,883Scored outfits
68.6Worn-look mean
57.4Pieces mean
11.2Point gap
How to read the numbers Every headline figure in this report is frozen on 2 October 2026. The page does not live-query production. If the corpus grows tomorrow, these n= values stay. Cite the snapshot, not "current site stats."

Section 1Methodology

The unit of analysis is one row in the OutfitScore analyses table. I keep a row if analysis_type is fashion or complete, analysis_result is present, and a numeric overall score can be read from score or overall_rating.score. Roast, makeup, body-type, color-season, and accessory-only analyses are excluded. Soft-deleted rows are not filtered: a deleted analysis is still a real observation of how the model labelled an upload.

The window is not a rolling 30-day panel. It is every qualifying row then in production: 27 September 2025 through 2 October 2026. That is 371 days. Report 06 used a later start and a male-only filter because it was a men's-style study. This report is not a gender study, so the window keeps every scored fashion and complete row.

Score extraction follows the same numeric-regex guard used by seo_stats_service. Non-numeric strings such as "N/A" become null and drop the row from the scored corpus. After that filter, n = 20,883. Mean 66.3, median 66.0, SD 15.1. Share scoring 85 or above: 12.8%. Share scoring 75 or above: 30.1%. Share below 60: 23.6%.

How an item-type label is defined

OutfitScore stores item_type as free text written by the model, not as a closed enum. I do not merge near-duplicates. complete_outfit (mean 68.6, n = 13,469) is a different string from complete_outfit_on_a_person (mean 68.8, n = 113). Merging them would invent a "worn look" family that the model did not write. The headline trio uses exact equality:

item_type = 'complete_outfit'
item_type = 'individual_pieces'
item_type = 'flat_lay'

A non-empty item_type is present on 14,955 of 20,883 scored rows (71.6%). Null or empty labels cover 5,928 rows (mean 63.6). Those rows are in the corpus totals only. They are not a type, and they are not in any ladder cell.

Publication floor A rate is printed only if the exact-label n is at least 100. clothing (n = 37, mean 0.0) and individual_piece (n = 9) exist in the corpus and are named as too thin to rate. Normalized buckets that collapse complete_outfit with complete_outfit_on_a_person are computed internally as a sensitivity check and are not published as headlines.

False labels are real. The model can write complete_outfit for a crop that stops at the chest, and individual_pieces for a rack of several garments. The method is exact string equality on model prose, not a human-coded content analysis with inter-rater kappa. Report 05 used regex on improvement tips. Report 06 used exact style_category. This report is the same method on a different field: three exact labels, no gender filter, no keyword hunt.

Section 2What this report is not

One OutfitScore post already occupies the pieces-versus-outfits shelf. Cloning it would be a second URL, not a second finding.

Already published What it is This report
Individual Pieces vs Full Outfits How-to on n = 86 vs 707; claims a 55–72 piece band Contrast only. Stale n=. Range claim is false here.
Report No. 06: The One-Notch Lift Men's exact style-category ladder Contrast only. Different field, mixed gender.
Exact complete_outfit vs individual_pieces vs flat_lay n = 13,469 / 957 / 218 Headline

The April post tells you when to upload a single garment and when to upload a full look. It does not publish a current n= for the three exact labels, and its 55–72 band does not survive this snapshot. That is this page. I do not rewrite the blog's n= here. That is a separate leftover.

This is also not a photo-type study. The photo_type column is null on 20,454 of the scored rows. Searching it would be a null hunt, not a type. Report 07 does not pretend otherwise.

Section 3The gender non-finding

Before the ladder: the comparison everyone will ask for. On exact complete_outfit, male rows mean 68.6 (n = 7,174) and female rows mean 68.2 (n = 5,816). On exact individual_pieces, male rows mean 58.0 (n = 453) and female rows mean 54.9 (n = 405). The worn-look gap is four tenths of a point. The piece gap is larger and still not the finding: both genders sit well below the worn-look band. Same refusal as Report 06. The story is the label, not the wearer.

Worn looks mean 68.6. Pieces mean 57.4. The story is not who uploaded. The story is whether the photograph had a body in it.

Section 4The context ladder

Four exact strings clear n ≥ 100. Three of them sit on a single context line: a worn look, a flat lay, a lone piece. The fourth, complete_outfit_on_a_person, is a near-duplicate of the worn-look string and is tabled without being merged.

Exact label n Mean Median 75+ 85+ <60 SD
complete_outfit 13,469 68.6 68.0 37.9% 16.8% 19.2% 14.5
complete_outfit_on_a_person 113 68.8 70.0 40.7% 14.2% 17.7% 13.2
flat_lay 218 64.1 63.5 18.8% 6.9% 28.9% 11.6
individual_pieces 957 57.4 61.0 22.4% 8.3% 45.4% 19.6
Mean score, exact item_type labels Published labels only where n ≥ 100. Dashed line = corpus mean 66.3. 90 75 60 45 30 66.3 68.6 Worn look n = 13,469 64.1 Flat lay n = 218 57.4 Pieces n = 957 Worn minus pieces is 11.2 points. Axis starts at 30 so the gaps are readable; it is not a zero baseline. Headline labels Corpus mean 66.3
Figure 2. Exact-label means on the worn–flat-lay–pieces line. Axis starts at 30. Presence of a label is not a treatment effect. complete_outfit_on_a_person (68.8, n = 113) is omitted from the chart so it is not merged into the worn-look bar.

The below-60 cut is louder than the mean. 45.4% of exact pieces sit below 60. 19.2% of worn looks do. 28.9% of flat lays do. The 75-plus cut runs the other way: 37.9% of worn looks, 22.4% of pieces, 18.8% of flat lays. Pieces can clear 75 — they do, 214 of 957 times — but the left tail is the story. p10 on pieces is 25. p10 on worn looks is 55.

Two mechanisms explain the ladder without requiring a treatment effect. First, the rubric itself is structural: colour harmony, styling cohesion, and occasion cannot be fully scored on a lone garment, so part of the gap is missing dimensions, not missing taste. Second, a look the model already likes is more likely to be photographed on a body. Those two facts are entangled. This snapshot cannot unentangle them.

Not a lift Worn-look rows mean 68.6. Piece rows mean 57.4. That is an 11.2-point observational gap between two labels, not evidence that uploading a full outfit raises a given garment by 11.2 points. The how-to for when to upload which already exists. This page will not restate it as a causal claim.

Section 5The range myth

The April blog said individual pieces cluster in a 55–72 band because three scoring dimensions cannot fire. On this snapshot that sentence is false.

Exact label Share in 55–72 p10 p90 SD Min Max
complete_outfit 51.5% 55.0 88.0 14.5 0 98
flat_lay 61.5% 49.0 80.0 11.6 25 89
individual_pieces 42.9% 25.0 80.0 19.6 0 95

Pieces are the widest of the three published labels, not the tightest. Only 42.9% of them sit in the blog's 55–72 band. Worn looks sit there more often (51.5%). Flat lays sit there most often (61.5%) and have the smallest SD (11.6). If any format is a mid-band cluster, it is the flat lay, not the lone piece.

A piece that scores 25 is usually not "three dimensions switched off." It is a garment the model already dislikes on quality, construction, or crop — or a zero from a failed parse that still passed the numeric regex as 0. A piece that scores 95 is a garment the model likes on the dimensions it can see. The missing-dimension story predicts a compressed middle. The data show a fat left tail.

Practical reading, not a lift If you are deciding whether to photograph a hanger or a body, the cheaper question is not "will the piece score in the 60s either way?" It is "do I want the left tail?" The corpus cannot tell you that a full-length shot will add eleven points to this jacket. It can tell you that exact pieces and exact worn looks do not live in the same band, and that the 55–72 cage was never locked.

Section 6Flat lay sits between them

Flat lay is the third published rung because it clears n ≥ 100 and because it is the photograph people take when they have a whole outfit and no body in the frame. Mean 64.1 sits 4.5 points below worn looks and 6.7 points above pieces. The below-60 rate (28.9%) sits between them too. The 75-plus rate (18.8%) is the one cell that does not interpolate: it is lower than pieces (22.4%). A laid-out look is less likely to crash, and less likely to clear 75, than a lone garment. That is a compression, and it is the opposite of the myth the blog attached to pieces.

complete_outfit_on_a_person (n = 113, mean 68.8) is printed so a reader can see that the near-duplicate of the worn-look string lands in the same band. Merging the two would add 113 rows and move the worn-look mean by less than a tenth of a point. I still do not merge them. Report 06 taught that lesson on streetwear punctuation. The method does not change because the means happen to agree.

Section 7What "context" means in the photograph

The how-to post already names the use cases: purchase decisions on a product shot, quality checks on a single garment, full-look feedback on a worn outfit. This report does not re-list them. It counts the label the model assigns after it has looked at the frame.

On this snapshot, moving from exact pieces to exact worn looks is associated with three changes in the score distribution at once: the mean steps 11.2 points, the below-60 rate halves, and the p10 jumps from 25 to 55. Moving from pieces to flat lay is a smaller mean step (6.7) with a tighter SD. That is why "context" is the right phrase and "twelve points" is the wrong one. The context is whether the model can see a body, a layout, or a lone object. The points are what the distribution did on labelled rows. They are not a promised increment for a reader who re-shoots tonight.

Section 8Limitations

The corpus is voluntary. People who upload an outfit to an AI rater are not a random sample of dressers. They are more likely to photograph a full look, more likely to be curious about a number, and more likely to own a garment worth a product shot. Hanger photos of fast-fashion returns that never get scored are under-counted by construction.

The classifier is exact string equality on model English. Punctuation and spacing split families that a human would merge. That split is treated as a feature, not a bug: publishing the merged worn-look number would hide that complete_outfit and complete_outfit_on_a_person are two strings even when their means agree. Singular individual_piece (n = 9) is missed by construction. I did not double-code a human gold set for this report.

Part of the gap is structural in the rubric. Colour harmony across garments, styling cohesion, and occasion-appropriateness cannot be fully evaluated on a lone item. A lower piece mean is therefore partly "the model refused to invent context," not "the garment was worse." This snapshot cannot say what share of the 11.2 points is that refusal.

Zeros exist. Exact pieces include scores of 0. Some of those are genuine lows; some may be parse leftovers that still matched the numeric regex. They pull the piece mean down and fatten the left tail. I do not drop zeros: a published 0 is still a row the model wrote. Sensitivity without zeros is not a headline.

Thin labels are named and withheld. Absence of a published clothing rate is not evidence that clothing-as-type does not happen. It is evidence that this snapshot cannot support a percentage without looking precise.

Section 9Conclusion

On 20,883 scored fashion and complete analyses, the mean is 66.3. The story is the exact-label ladder on item_type: complete_outfit 68.6 (n = 13,469), flat_lay 64.1 (n = 218), individual_pieces 57.4 (n = 957). Worn minus pieces is 11.2 points. The below-60 rate goes 19.2% → 28.9% → 45.4%. Pieces are not a 55–72 band: 42.9% land there, SD is 19.6, p10 is 25. Labels are not merged. n below 100 is not printed as a rate. No causal lift is claimed.

The honest sentence for a journalist is not "upload a full outfit and gain eleven points." It is: when this model labels an upload a complete outfit rather than individual pieces, the two groups do not share a band. If you want the engine to look at your own frame, the tool is the same as it was: the free AI outfit rater. This page is the study. Pitch this URL.

How To Cite This Report

Title: The Context Gap: What 20,883 Outfit Scores Reveal About Worn Looks vs Individual Pieces
Author: Saad, Founder of OutfitScore
Publication: OutfitScore Research Reports, No. 07
Date: October 2, 2026
Sample: n = 20,883 scored fashion/complete analyses, 27 Sep 2025 – 2 Oct 2026
Saad. (2026). The Context Gap: What 20,883 Outfit Scores Reveal About Worn Looks vs Individual Pieces. OutfitScore Research Reports, No. 07. Retrieved from https://outfitscore.com/research/the-context-gap