The One-Notch Lift
What 10,050 men's outfit scores reveal about Casual vs Smart Casual
style_category strings
that clear n ≥ 100, Smart Casual means 76.4 (n = 433) against
Casual 64.5 (n = 1,104) and Casual/Loungewear 52.6 (n = 112).
Each step is 11.9 points. The 75-plus rate is the sharper cut: 61.7% of exact Smart Casual
versus 20.2% of Casual versus 0.9% of Casual/Loungewear. Labels are exact string matches,
not merged families. This is an observational snapshot of model labels, not a treatment
study. No causal score lift is claimed.
Men's style media will sell you ten essentials, a capsule, or a blazer-and-loafer swap. OutfitScore already published two of those how-tos. This report is the n= behind the swap: what the model actually writes when the wearer is male and the label is Casual, Smart Casual, or Loungewear.
The tempting headline is that men score worse than women. They do not. On this snapshot the
male mean is 67.1 and the female mean is 66.5 — six tenths of a point, on samples of 10,050
and 9,128. That gap is not a finding. It is a refusal. The finding sits inside the men's
corpus, on the exact strings the model uses for style_category.
Section 1Methodology
The unit of analysis is one row in the OutfitScore analyses table. I keep a row if
analysis_type is fashion or complete,
lower(analysis_result->>'gender') equals male,
analysis_result is present, and a numeric overall score can be read from
score or overall_rating.score. Roast, makeup, body-type, color-season,
and accessory-only analyses are excluded. Soft-deleted rows are not filtered: a deleted analysis
is still a real observation of how the model labelled an outfit.
The window is not a rolling 30-day panel. It is every qualifying row then in production: 1 November 2025 through 2 October 2026. That is 336 days, not a season. Seasonal mix (shorts vs wool, boots vs sneakers) is therefore averaged, not isolated.
Score extraction follows the same numeric-regex guard used by seo_stats_service.
Non-numeric strings such as "N/A" become null and drop the row from the scored corpus. After that
filter, n = 10,050. Mean 67.1, median 66.0. Share scoring 85 or above: 12.2%. Share scoring 75 or
above: 30.7%. Share below 60: 21.5%.
How a style label is defined
OutfitScore stores style_category as free text written by the model, not as a closed
enum. I do not merge near-duplicates. "Casual / Streetwear" (mean 76.7, n = 319) is a different
string from "Streetwear" (mean 65.0, n = 454). Merging them would invent a family that the model
did not write. The headline trio uses exact equality:
style_category = 'Casual' style_category = 'Smart Casual' style_category = 'Casual/Loungewear'
A non-empty style_category is present on 7,985 of 10,050 scored male rows (79.5%).
The remaining 20.5% have no label and are in the corpus totals only. They are not in any
ladder cell.
False labels are real. The model can write "Smart Casual" for a look that a human rater would call business-casual, and "Casual" for a look that is already a blazer away from the next register. The method is exact string equality on model prose, not a human-coded content analysis with inter-rater kappa. Report 05 used regex on improvement tips. This report is narrower: three exact labels, one gender filter, no keyword hunt.
Section 2What this report is not
Two OutfitScore posts already occupy the men's-style and smart-casual shelf. Cloning them would be a second URL, not a second finding.
| Already published | What it is | This report |
|---|---|---|
| Men's Style Essentials | How-to wardrobe list | Contrast only |
| Casual to Smart Casual, One Swap | Blazer-and-loafer how-to | Contrast only |
| Smart Casual dress code | Definition page | Contrast only |
| Exact Casual vs Smart Casual vs Loungewear, men only | n = 1,104 / 433 / 112 | Headline |
The essentials guide tells you which ten pieces to own. The one-swap post tells you which garment to add. The dress-code page tells you what the room asked for. None of them publishes a men's n= for the three exact labels. That is this page.
This is also not a body-type study. Fashion rows do not store a structured body-type field. Searching the JSON for "pear" or "hourglass" is a word hunt, not a type. Report 06 does not pretend otherwise.
Section 3The gender non-finding
Before the ladder: the comparison everyone will ask for.
| Corpus | n scored | Mean | Median |
|---|---|---|---|
Male (gender = 'male') |
10,050 | 67.1 | 66.0 |
Female (gender = 'female') |
9,128 | 66.5 | 66.0 |
The mean gap is 0.6 points. The medians are identical. Fashion media that frames "men dress worse" is not looking at this corpus. The honest sentence is that men and women who upload an outfit to this rater land in the same band. The interesting variance is inside the men's labels, not between genders.
The men's mean is 67.1. The women's mean is 66.5. The story is not that gap. The story is the 11.9-point step inside the men's labels.
Section 4The one-notch ladder
Three exact strings clear n ≥ 100 and sit on a single formality line: loungewear, casual, smart casual. Everything else in the labelled set is a side path (streetwear, athleisure, minimalist) and is tabled in Section 5 without being merged into these three.
| Exact label | n | Mean | Median | 75+ | 85+ | <60 |
|---|---|---|---|---|---|---|
| Casual/Loungewear | 112 | 52.6 | 53.0 | 0.9% | 0.0% | 82.1% |
| Casual | 1,104 | 64.5 | 65.0 | 20.2% | 4.8% | 24.7% |
| Smart Casual | 433 | 76.4 | 76.0 | 61.7% | 28.2% | 2.8% |
The 75-plus cut is louder than the mean. Six in ten exact Smart Casual rows clear 75. Two in ten Casual rows do. Almost none of the Loungewear rows do (1 of 112). The below-60 cut runs the other way: 82.1% of Loungewear, 24.7% of Casual, 2.8% of Smart Casual.
Two mechanisms explain the ladder without requiring a treatment effect. First, the label is partly the score: a look the model already likes is more likely to be called Smart Casual. Second, Smart Casual as a garment stack — structured layer, cleaner shoe, visible break — is the same stack the rubric rewards on fit, cohesion, and occasion. Those two facts are entangled. This snapshot cannot unentangle them.
Section 5The rest of the labelled set
Thirteen exact strings clear n ≥ 100. The headline trio is three of them. The other ten are published so a reader can see that "streetwear" is not one number, and that merging would have been a lie.
Exact style_category |
n | Mean | Median | 75+ | 85+ | <60 |
|---|---|---|---|---|---|---|
| Casual | 1,104 | 64.5 | 65.0 | 20.2% | 4.8% | 24.7% |
| Streetwear | 454 | 65.0 | 65.0 | 15.0% | 1.3% | 22.7% |
| Smart Casual | 433 | 76.4 | 76.0 | 61.7% | 28.2% | 2.8% |
| Casual / Streetwear | 319 | 76.7 | 78.0 | 68.0% | 33.5% | 4.4% |
| Casual Streetwear | 303 | 70.8 | 68.0 | 36.0% | 20.1% | 9.9% |
| Casual Minimalist | 248 | 66.4 | 65.0 | 11.7% | 2.8% | 8.5% |
| Casual Everyday | 208 | 77.1 | 76.5 | 73.1% | 29.8% | 3.8% |
| Casual/Streetwear | 172 | 62.3 | 63.5 | 7.0% | 0.0% | 27.9% |
| Athleisure | 162 | 60.4 | 61.0 | 11.7% | 2.5% | 43.2% |
| Casual / Athleisure | 151 | 72.7 | 75.0 | 57.0% | 23.2% | 11.9% |
| Streetwear / Casual | 127 | 79.3 | 80.0 | 76.4% | 43.3% | 1.6% |
| Casual/Loungewear | 112 | 52.6 | 53.0 | 0.9% | 0.0% | 82.1% |
| Streetwear/Casual | 102 | 66.4 | 65.0 | 15.7% | 2.9% | 16.7% |
Read the streetwear cluster before merging anything. Exact "Streetwear" means 65.0. "Casual / Streetwear" means 76.7. "Casual/Streetwear" (no spaces) means 62.3. "Streetwear / Casual" means 79.3. Those are four strings, four means, four sample sizes. A researcher who folds them into one "streetwear family" publishes a number that none of the rows actually has.
Casual Everyday (n = 208, mean 77.1) outscores exact Casual (n = 1,104, mean 64.5) by 12.6 points — a larger gap than the headline step — and is still not the headline, because "everyday" here is a style adjective, not the occasion field, and because the n is a fifth of Casual. Athleisure (n = 162, mean 60.4) sits closer to loungewear than to smart casual, which is consistent with the model's habit of scoring gym-adjacent fabric as unfinished for a street photograph.
Section 6Occasion is not the story either
Occasion cells that clear n ≥ 100 in the male scored corpus: everyday 4,780 (mean 66.0); general 2,754 (70.2); unspecified 1,316 (68.3); general style 260 (62.1); streetwear as occasion 154 (66.0). Everyday is half the men's volume and sits one point below the male mean. That is a mix note, not a finding. Date night, office, and wedding-guest cells exist and are below the publication floor on this snapshot. They are named so their absence is not mistaken for "men do not dress for those rooms."
The label ladder in Section 4 is therefore not an occasion ranking in disguise. Most of the Casual and Smart Casual rows are everyday or general. The 11.9-point step is a register step inside the same rooms, not a comparison of black-tie against school run.
Section 7What "one notch" means in the photograph
The how-to posts already name the stack: a structured layer, a leather or suede shoe, a visible break at the ankle. Report 05 counted the hem as an unpublished finish flag on the mixed-gender corpus. This report does not re-count hems. It counts the label the model assigns after it has looked at the whole stack.
On this snapshot, moving from exact Casual to exact Smart Casual is associated with three changes in the score distribution at once: the mean steps 11.9 points, the 75-plus rate triples, and the below-60 rate falls from one in four to one in thirty-six. Moving from Loungewear to Casual is the same mean step in the other direction, with an even harsher below-60 collapse (82.1% to 24.7%).
That is why "one notch" is the right phrase and "twelve points" is the wrong one. The notch is a register: out of the house-clothes label, or out of the default-casual label. The points are what the distribution did on labelled rows. They are not a promised increment for a reader who buys a blazer tonight.
Section 8Limitations
The corpus is voluntary. People who upload an outfit to an AI rater are not a random sample of men. They are more likely to photograph a full look, more likely to be curious about a number, and more likely to own the kind of garment that can be labelled Smart Casual — a blazer, a loafer, a chino. Tracksuits, uniforms, and workwear that never get photographed for a score are under-counted by construction.
Gender is a model field, not a self-report. Rows labelled unclear (591),
unisex (141), and other leftovers are excluded from both the male corpus
and the female contrast. A different gender taxonomy would move rows.
The classifier is exact string equality on model English. Punctuation and spacing split families that a human would merge. That split is treated as a feature in Section 5, not a bug: publishing the merged number would hide that "Casual / Streetwear" and "Casual/Streetwear" do not mean the same score. Multilingual labels and novel phrasings are missed. I did not double-code a human gold set for this report.
Photo crop is unmeasured here. A mirror selfie that stops at the chest cannot support a Smart Casual call that depends on the shoe. If Smart Casual rows are more often full-length, part of the 11.9-point gap is a photograph effect. Report-style photo-type work exists elsewhere on the site; it is not merged into these denominators.
Thin labels are named and withheld. Absence of a published Business Casual rate is not evidence that the code does not matter. It is evidence that this snapshot cannot support a percentage without looking precise.
Section 9Conclusion
On 10,050 scored men's fashion and complete analyses, the mean is 67.1 — six tenths above the parallel female corpus, which is not a story. The story is the exact-label ladder inside the men's set: Casual/Loungewear 52.6 (n = 112), Casual 64.5 (n = 1,104), Smart Casual 76.4 (n = 433). Each step is 11.9 points. The 75-plus rate goes 0.9% → 20.2% → 61.7%. Labels are not merged. n below 100 is not printed as a rate. No causal lift is claimed.
The honest sentence for a journalist is not "dressing one notch raises your score by twelve points." It is: when this model labels a man's outfit Smart Casual rather than Casual, the two groups do not share a band. If you want the engine to look at your own register, the tool is the same as it was: the free AI outfit rater. This page is the study. Pitch this URL.