The Context Gap
What 20,883 outfit scores reveal about worn looks vs individual pieces
item_type strings
that clear n ≥ 100, complete_outfit means 68.6 (n = 13,469) against
individual_pieces 57.4 (n = 957) and flat_lay 64.1 (n = 218).
The worn-versus-pieces gap is 11.2 points. The below-60 rate is the sharper cut:
45.4% of exact pieces versus 19.2% of worn looks. Pieces are not a tight 55–72 band:
only 42.9% land there, and the piece SD is 19.6 against 14.5 for worn looks. Labels are
exact string matches, not merged families. This is an observational snapshot of model
labels, not a treatment study. No causal score lift is claimed.
OutfitScore already published a how-to on individual pieces versus full outfits.
That post used 86 piece analyses and 707 complete looks, and it claimed pieces cluster in a
55–72 band. This report is the n= behind that shelf: what the model actually writes when the
label is complete_outfit, individual_pieces, or flat_lay,
on a corpus more than twenty times larger.
The tempting headline is that a lone garment cannot score high. It can. Exact pieces reach 95. The honest headline is that they also reach 0, that nearly half of them sit below 60, and that a worn look and a piece do not share a band. The April blog's tight mid-band is false on this snapshot. The how-to stays a how-to. This page is the study.
Section 1Methodology
The unit of analysis is one row in the OutfitScore analyses table. I keep a row if
analysis_type is fashion or complete,
analysis_result is present, and a numeric overall score can be read from
score or overall_rating.score. Roast, makeup, body-type, color-season,
and accessory-only analyses are excluded. Soft-deleted rows are not filtered: a deleted analysis
is still a real observation of how the model labelled an upload.
The window is not a rolling 30-day panel. It is every qualifying row then in production: 27 September 2025 through 2 October 2026. That is 371 days. Report 06 used a later start and a male-only filter because it was a men's-style study. This report is not a gender study, so the window keeps every scored fashion and complete row.
Score extraction follows the same numeric-regex guard used by seo_stats_service.
Non-numeric strings such as "N/A" become null and drop the row from the scored corpus. After that
filter, n = 20,883. Mean 66.3, median 66.0, SD 15.1. Share scoring 85 or above: 12.8%. Share
scoring 75 or above: 30.1%. Share below 60: 23.6%.
How an item-type label is defined
OutfitScore stores item_type as free text written by the model, not as a closed
enum. I do not merge near-duplicates. complete_outfit (mean 68.6, n = 13,469) is a
different string from complete_outfit_on_a_person (mean 68.8, n = 113). Merging them
would invent a "worn look" family that the model did not write. The headline trio uses exact
equality:
item_type = 'complete_outfit' item_type = 'individual_pieces' item_type = 'flat_lay'
A non-empty item_type is present on 14,955 of 20,883 scored rows (71.6%).
Null or empty labels cover 5,928 rows (mean 63.6). Those rows are in the corpus totals only.
They are not a type, and they are not in any ladder cell.
clothing (n = 37, mean 0.0)
and individual_piece (n = 9) exist in the corpus and are named as too thin to rate.
Normalized buckets that collapse complete_outfit with
complete_outfit_on_a_person are computed internally as a sensitivity check and are
not published as headlines.
False labels are real. The model can write complete_outfit for a crop that stops at
the chest, and individual_pieces for a rack of several garments. The method is exact
string equality on model prose, not a human-coded content analysis with inter-rater kappa.
Report 05 used regex on improvement tips. Report 06 used exact style_category.
This report is the same method on a different field: three exact labels, no gender filter, no
keyword hunt.
Section 2What this report is not
One OutfitScore post already occupies the pieces-versus-outfits shelf. Cloning it would be a second URL, not a second finding.
| Already published | What it is | This report |
|---|---|---|
| Individual Pieces vs Full Outfits | How-to on n = 86 vs 707; claims a 55–72 piece band | Contrast only. Stale n=. Range claim is false here. |
| Report No. 06: The One-Notch Lift | Men's exact style-category ladder | Contrast only. Different field, mixed gender. |
| Exact complete_outfit vs individual_pieces vs flat_lay | n = 13,469 / 957 / 218 | Headline |
The April post tells you when to upload a single garment and when to upload a full look. It does not publish a current n= for the three exact labels, and its 55–72 band does not survive this snapshot. That is this page. I do not rewrite the blog's n= here. That is a separate leftover.
This is also not a photo-type study. The photo_type column is null on 20,454 of
the scored rows. Searching it would be a null hunt, not a type. Report 07 does not pretend
otherwise.
Section 3The gender non-finding
Before the ladder: the comparison everyone will ask for. On exact complete_outfit,
male rows mean 68.6 (n = 7,174) and female rows mean 68.2 (n = 5,816). On exact
individual_pieces, male rows mean 58.0 (n = 453) and female rows mean 54.9
(n = 405). The worn-look gap is four tenths of a point. The piece gap is larger and still
not the finding: both genders sit well below the worn-look band. Same refusal as Report 06.
The story is the label, not the wearer.
Worn looks mean 68.6. Pieces mean 57.4. The story is not who uploaded. The story is whether the photograph had a body in it.
Section 4The context ladder
Four exact strings clear n ≥ 100. Three of them sit on a single context line: a worn look,
a flat lay, a lone piece. The fourth, complete_outfit_on_a_person, is a near-duplicate
of the worn-look string and is tabled without being merged.
| Exact label | n | Mean | Median | 75+ | 85+ | <60 | SD |
|---|---|---|---|---|---|---|---|
| complete_outfit | 13,469 | 68.6 | 68.0 | 37.9% | 16.8% | 19.2% | 14.5 |
| complete_outfit_on_a_person | 113 | 68.8 | 70.0 | 40.7% | 14.2% | 17.7% | 13.2 |
| flat_lay | 218 | 64.1 | 63.5 | 18.8% | 6.9% | 28.9% | 11.6 |
| individual_pieces | 957 | 57.4 | 61.0 | 22.4% | 8.3% | 45.4% | 19.6 |
complete_outfit_on_a_person (68.8, n = 113) is omitted from the chart so it is not merged into the worn-look bar.The below-60 cut is louder than the mean. 45.4% of exact pieces sit below 60. 19.2% of worn looks do. 28.9% of flat lays do. The 75-plus cut runs the other way: 37.9% of worn looks, 22.4% of pieces, 18.8% of flat lays. Pieces can clear 75 — they do, 214 of 957 times — but the left tail is the story. p10 on pieces is 25. p10 on worn looks is 55.
Two mechanisms explain the ladder without requiring a treatment effect. First, the rubric itself is structural: colour harmony, styling cohesion, and occasion cannot be fully scored on a lone garment, so part of the gap is missing dimensions, not missing taste. Second, a look the model already likes is more likely to be photographed on a body. Those two facts are entangled. This snapshot cannot unentangle them.
Section 5The range myth
The April blog said individual pieces cluster in a 55–72 band because three scoring dimensions cannot fire. On this snapshot that sentence is false.
| Exact label | Share in 55–72 | p10 | p90 | SD | Min | Max |
|---|---|---|---|---|---|---|
| complete_outfit | 51.5% | 55.0 | 88.0 | 14.5 | 0 | 98 |
| flat_lay | 61.5% | 49.0 | 80.0 | 11.6 | 25 | 89 |
| individual_pieces | 42.9% | 25.0 | 80.0 | 19.6 | 0 | 95 |
Pieces are the widest of the three published labels, not the tightest. Only 42.9% of them sit in the blog's 55–72 band. Worn looks sit there more often (51.5%). Flat lays sit there most often (61.5%) and have the smallest SD (11.6). If any format is a mid-band cluster, it is the flat lay, not the lone piece.
A piece that scores 25 is usually not "three dimensions switched off." It is a garment the model already dislikes on quality, construction, or crop — or a zero from a failed parse that still passed the numeric regex as 0. A piece that scores 95 is a garment the model likes on the dimensions it can see. The missing-dimension story predicts a compressed middle. The data show a fat left tail.
Section 6Flat lay sits between them
Flat lay is the third published rung because it clears n ≥ 100 and because it is the photograph people take when they have a whole outfit and no body in the frame. Mean 64.1 sits 4.5 points below worn looks and 6.7 points above pieces. The below-60 rate (28.9%) sits between them too. The 75-plus rate (18.8%) is the one cell that does not interpolate: it is lower than pieces (22.4%). A laid-out look is less likely to crash, and less likely to clear 75, than a lone garment. That is a compression, and it is the opposite of the myth the blog attached to pieces.
complete_outfit_on_a_person (n = 113, mean 68.8) is printed so a reader can
see that the near-duplicate of the worn-look string lands in the same band. Merging the
two would add 113 rows and move the worn-look mean by less than a tenth of a point. I
still do not merge them. Report 06 taught that lesson on streetwear punctuation. The
method does not change because the means happen to agree.
Section 7What "context" means in the photograph
The how-to post already names the use cases: purchase decisions on a product shot, quality checks on a single garment, full-look feedback on a worn outfit. This report does not re-list them. It counts the label the model assigns after it has looked at the frame.
On this snapshot, moving from exact pieces to exact worn looks is associated with three changes in the score distribution at once: the mean steps 11.2 points, the below-60 rate halves, and the p10 jumps from 25 to 55. Moving from pieces to flat lay is a smaller mean step (6.7) with a tighter SD. That is why "context" is the right phrase and "twelve points" is the wrong one. The context is whether the model can see a body, a layout, or a lone object. The points are what the distribution did on labelled rows. They are not a promised increment for a reader who re-shoots tonight.
Section 8Limitations
The corpus is voluntary. People who upload an outfit to an AI rater are not a random sample of dressers. They are more likely to photograph a full look, more likely to be curious about a number, and more likely to own a garment worth a product shot. Hanger photos of fast-fashion returns that never get scored are under-counted by construction.
The classifier is exact string equality on model English. Punctuation and spacing split
families that a human would merge. That split is treated as a feature, not a bug:
publishing the merged worn-look number would hide that complete_outfit and
complete_outfit_on_a_person are two strings even when their means agree.
Singular individual_piece (n = 9) is missed by construction. I did not
double-code a human gold set for this report.
Part of the gap is structural in the rubric. Colour harmony across garments, styling cohesion, and occasion-appropriateness cannot be fully evaluated on a lone item. A lower piece mean is therefore partly "the model refused to invent context," not "the garment was worse." This snapshot cannot say what share of the 11.2 points is that refusal.
Zeros exist. Exact pieces include scores of 0. Some of those are genuine lows; some may be parse leftovers that still matched the numeric regex. They pull the piece mean down and fatten the left tail. I do not drop zeros: a published 0 is still a row the model wrote. Sensitivity without zeros is not a headline.
Thin labels are named and withheld. Absence of a published clothing rate is
not evidence that clothing-as-type does not happen. It is evidence that this snapshot
cannot support a percentage without looking precise.
Section 9Conclusion
On 20,883 scored fashion and complete analyses, the mean is 66.3. The story is the
exact-label ladder on item_type: complete_outfit 68.6 (n = 13,469),
flat_lay 64.1 (n = 218), individual_pieces 57.4 (n = 957). Worn minus pieces is 11.2
points. The below-60 rate goes 19.2% → 28.9% → 45.4%. Pieces are not a 55–72 band:
42.9% land there, SD is 19.6, p10 is 25. Labels are not merged. n below 100 is not
printed as a rate. No causal lift is claimed.
The honest sentence for a journalist is not "upload a full outfit and gain eleven points." It is: when this model labels an upload a complete outfit rather than individual pieces, the two groups do not share a band. If you want the engine to look at your own frame, the tool is the same as it was: the free AI outfit rater. This page is the study. Pitch this URL.