Design research One brief, five backends, 182 sources. Four returned; one ran 165 minutes and was abandoned.

Dossier · Email design · 21 August 2026

The item count was never the defect

A digest carrying twenty-four items went out and was judged unreadable. Four independent research backends were given the same brief. Three of them recommended cutting the list. The two best-sourced showed there is no evidence for an item ceiling at all, and named a different defect entirely.

Read this as

The short version

Twenty-four things in one email is not what made it unreadable. All twenty-four looking equally important is.

The defect in a flat twenty-four-item digest is undifferentiated scan cost per item, not the number of items.

The defect is undifferentiated scan cost per item: n identical evaluations inside a fixed attention budget, with no visual signal permitting early exit.

What would change it Nobody has ever run the actual test. If someone sent the same twenty-four items as a flat list to half a list and as tiers to the other half, and the tiers did no better, this whole page falls over. No published study compares a tiered multi-item email against a flat list of the same items. All four backends searched for one. A within-programme randomised test measuring read time and per-tier click penetration would settle it, and could overturn the central recommendation. The tiering recommendation is an inference from four measured results, not a measured comparison. A within-programme randomised split holding subject, content and item order constant, measured on bot-filtered unique clicks per delivered email and per-tier penetration, is the test that would confirm or overturn it. Four of the seven questions in the brief have no published study behind them at all.

01; The negative finding

Cutting the list was the wrong fix, and three of four backends recommended it

When an email with twenty-four things in it goes badly, the obvious response is to put fewer things in it. That was the recommendation from three of the four research systems. It is also the one that the evidence does not support.

The intuitive fix is fewer items. Three of the four backends recommended some version of it, one of them citing "extreme click decay" in a flat list. The two best-sourced members went the other way, and the disagreement turns out to be about where the numbers came from.

The intuitive fix is a lower item count, and it was recommended by three of four members. It does not survive source-checking. The decisive detail is the provenance of the decay figures rather than their magnitude.

Direct finding High confidence

The biggest piece of evidence anyone has on this looked at 317,000 email campaigns and 2.9 billion emails. Emails with more than twenty links got more clicks from the people who opened them than emails with fewer.1MailerLite21+ links: highest CTOR in the dataset at 6.72%.12 February 2026

MailerLite's analysis of 317,000 campaigns and 2.9 billion emails found the 21-or-more-links bucket carried the highest click-to-open rate in the dataset at 6.72%. Campaign Monitor, testing the same hypothesis independently on its own data, found click rate rising with unique link count to around eleven links and then plateauing.1MailerLite317,000 campaigns / 2.9bn emails. 21+ links: highest CTOR 6.72%, lowest open rate 29.9%.12 February 20265Campaign Monitor, Do Fewer Links Mean More Clicks?Found the opposite of what it expected: more links correlate with higher click rates, rising to ~11 then holding.2019 · observational

MailerLite, n=317,000 campaigns / 2.9bn emails, counting unique URLs with unsubscribe, header and footer links excluded: the 21+ bucket carries the highest CTOR at 6.72%, described as 16.67% above their average, alongside the lowest open rate at 29.9%. Campaign Monitor's independent test on its own campaign data found CTR rising to approximately eleven unique links then holding near 17%.1MailerLiteObservational, not controlled. Link count correlates with sender category rather than being randomly assigned.12 February 20265Campaign MonitorObservational campaign data, 2019 update.2019

LimitsBoth are observational campaign data rather than controlled experiments, and link count correlates with sender category rather than being randomly assigned. MailerLite's own headline recommendation is 2 to 5 links, driven by its open-rate and e-commerce conversion columns, which are the two metrics least applicable to a non-commerce product digest. The two datasets agree on the click-side direction and diverge only where those columns enter the ranking.

  • 1 link5.9%↓
  • 2–5 links6.08%
  • 21+ links6.72%
Click-to-open rate by unique link count. The bucket that should have collapsed is the one that leads. Among people who opened, more items meant more clicking. Chart shows CTOR only; the same dataset's open-rate column runs the other way, and is the column Apple's Mail Privacy Protection broke. Source: MailerLite, 317,000 campaigns and 2.9 billion emails, February 2026.1MailerLiteUnique URLs only, with unsubscribe, header and footer links excluded.12 February 2026
Direct finding Medium confidence

The theory people reach for is that too much choice overwhelms you. Researchers pooled every test of that idea they could find. The effect came out at roughly nothing.2Scheibehenne et al., JCR 37(3)63 conditions, 50 experiments, N=5,036. Mean effect size virtually zero.2010

The theoretical backstop for an item cap is choice overload, and the pooled evidence does not support it. Scheibehenne, Greifeneder and Todd's meta-analysis of 63 conditions across 50 experiments found a mean effect size of virtually zero, and could not identify sufficient conditions for the effect.2Scheibehenne, Greifeneder & Todd, JCR 37(3) 409-425Meta-analysis, 63 conditions / 50 experiments / N=5,036.2010 · meta-analysis

Scheibehenne et al. 2010, Journal of Consumer Research 37(3) 409–425: mean effect size approximately zero across 63 conditions, 50 experiments, N=5,036, with the authors reporting they "could not identify sufficient conditions or specific circumstances" for choice overload.2Scheibehenne, Greifeneder & ToddJournal of Consumer Research 37(3), pp. 409-425.2010

LimitsChernev, Böckenholt and Goodman's competing meta-analysis of 99 experiments argues the effect does exist under specific moderators. The literature is genuinely contested rather than settled. What both exclude is a general numeric item threshold for a digest.3Chernev, Böckenholt & Goodman, JCPCompeting meta-analysis, 99 experiments across 53 papers, arguing for moderator-dependent effects.2015

Where the decay numbers actually came from

One backend argued that a flat list triggers extreme click decay, and gave numbers: the top item takes about 39.8% of clicks, the second 18.7%, the third 10.2%. Those figures are real. They are also measurements of search engine results pages, applied to email by analogy, and the backend that supplied them says so in its own text.

The decay figures offered in support of an item cap (position one 39.6–39.8%, position two 18.7%, position three 10.2%) are search-engine result CTR, transferred to email by explicit inference. The member supplying them tagged the transfer itself, and separately recorded that item-by-item decay data for a 24-item email digest does not exist in published form. The email-native equivalent is weaker and points elsewhere: Kumar and Salo's peer-reviewed analysis found newsletter click-through follows a U-pattern rather than a monotonic decay, with left-region links outperforming right.6Kumar & Salo, Journal of Marketing Communications 24(5) 535-548Peer-reviewed, LianaMailer analytics. Click-through follows a U-pattern, explicitly contrasted with the Z-pattern of web design.16 March 2016 · peer-reviewed

02; The mechanism

Fifty-one seconds buys rejections, not readings

People give a newsletter about fifty seconds. In that time they are not reading it. They are deciding, item by item, what to ignore.

If the item count is not the constraint, something else is. It is the attention budget, and the budget is spent on rejection rather than reading. That reframes the design question entirely: not how many items can be read, but how many can be rejected per second.

The binding constraint is a fixed attention budget spent predominantly on triage. The design target is rejection throughput, not reading throughput, and that is what makes uniform visual weight expensive.

Direct finding High confidence

Someone put forty-two people in front of a hundred and seventeen newsletters and tracked their eyes. Average time spent: fifty-one seconds. Properly read: about one in five. What they looked at was the first two words of each heading.4Nielsen Norman GroupEyetracking, n=42, 117 newsletters.11 June 2006

Nielsen Norman Group's eyetracking work measured average time-after-open at 51 seconds, with participants fully reading 19% of newsletters and glancing at or partially skimming 35%. Heatmaps concentrated on the first two words of the headlines.4Nielsen Norman Group, Email Newsletters: Surviving Inbox Congestionn=42, 117 newsletters. 51s mean, 19% fully read, 35% glanced or skimmed.11 June 2006 · eyetracking

NN/g eyetracking, n=42 participants across 117 newsletters: mean time-after-open 51 seconds; 19% fully read; 35% glanced at or partially skimmed; heatmap concentration on the first two words of headlines. This is the single strongest measured constraint in the corpus and the one that most directly justifies tiering.4Nielsen Norman GroupThe only eyetracking measurement of email newsletters located across 182 sources.11 June 2006

LimitsPublished in 2006. The measurement predates mobile email entirely, and no equivalent modern eyetracking study of email was located by any of the four backends. It is load-bearing here and it is nineteen years old, which is a real weakness in the argument rather than a footnote.

Inference High confidence

Put those together. All twenty-four items looked the same, so each one cost the same effort to judge, and nothing on the page said where to stop. That is the defect, and it does not go away by having fewer items. It goes away by making some items obviously bigger than others.4Nielsen Norman Group51s budget, 19% fully read.11 June 2006

In a flat list, every item costs the same fixation to evaluate, and nothing in the visual field tells the reader where to stop. Twenty-four identical evaluations do not fit in fifty-one seconds. A tier is a pre-made decision: it tells the reader, before they read anything, how much attention this item is claiming.4Nielsen Norman Group51s mean time-after-open; heatmaps on the first two words of headlines.11 June 2006

Composing the scan budget with the null item-count result: a flat list requires n identical evaluations inside a fixed budget with no visual signal permitting early exit, so evaluation cost scales linearly in n while the budget is constant. Tiering moves that decision to build time. This is assembled from four measured results rather than measured directly, and it is marked as an inference for that reason.

LimitsNo published study compares a tiered multi-item email against a flat list of the same items. All four backends searched Litmus, Email on Acid, NN/g and the academic literature and none found one. This is the central recommendation of the page and it rests on inference, which is why it should be evaluated in the generating skill's own evals rather than asserted as measured fact.

03; The only causal evidence

Prominence earns its space; ordering the tail does not

Across 182 sources, exactly one proper experiment exists. It ran for eight weeks and it tested the thing we care about.

One member found the only causal result in the entire corpus: a peer-reviewed eight-week field experiment that manipulated which items were featured. It supports the featured tier directly, and it kills a feature you might otherwise build.

Kong et al. 2022 is the sole causal study located across four independent literature sweeps. It bounds where design effort returns anything.

Direct finding High confidence

Putting the items a reader actually cared about at the top nearly doubled how many of them got read properly, from 13% to 22%.7Kong et al., ACM CSCWEight-week field experiment. Detail-reading 13%→22%, recognition 37%→49%.11 November 2022

Kong et al.'s eight-week field experiment, with 117 completed participant records and message-level data for 4,242 messages, found that placing reader-preferred messages in the top-news area raised their recognition from 37% to 49% and their detail-reading from 13% to 22%.7Kong et al., ACM CSCW117 completed participant records, recognition and detail-reading data for 4,242 individual messages.11 November 2022 · peer-reviewed

Kong et al. 2022, ACM CSCW, eight-week field experiment, 117 completed participant records, recognition and detail-reading recorded for 4,242 individual messages. Top-news placement of reader-preferred items: recognition 37%→49% (+12pp), detail-reading 13%→22% (+9pp). Mixing reader-preferred and organisation-preferred items in top news raised whole-newsletter recognition by 19pp versus randomised top news.7Kong et al., ACM CSCWThe strongest causal evidence located. Setting is an internal organisational newsletter.11 November 2022

LimitsThe setting is an internal university newsletter, not a public developer product digest. Transferability is an assumption rather than a demonstrated property, and the effect is conditional on the featured items being genuinely relevant to that reader, which a build-time tier cannot guarantee without personalisation.

Direct finding Medium confidence

The same experiment shuffled the order of everything below the featured block. It changed nothing at all.7Kong et al.Message-order treatments below top-news: no significant effects.11 November 2022

In the same experiment, changing the order of messages below the top-news area produced no significant effect on interest, reading time or overall recognition. The entire measured benefit sits in selective prominence, not in sequence.7Kong et al.Order treatments below the top-news area were not significant on interest, reading time or overall recognition.11 November 2022

Message-order treatments below the top-news area returned no significant effects on interest, reading time or overall recognition. Combined with the featuring result, this localises the whole measured benefit to selective prominence. The practical consequence is that a relevance-ranking system for the long tail has no evidential support, while feature selection does.7Kong et al.A null result within the same experiment that produced the positive featuring result.11 November 2022

LimitsA null result in one experiment in one setting. It bounds where effort is worth spending rather than proving that order never matters anywhere.

04; The block at the top

The summary survives, but not as a paragraph and not as a contents list

A summary at the top sounds obviously helpful. The evidence is more specific than that, and it rules out two of the three obvious ways to build one.

This was the sharpest disagreement in the panel. One member argued a contents block actively suppresses engagement. Another found no measured study either way. They converge anyway, because the measured evidence is about a narrower thing than either framing.

The satisfy-versus-amplify question is unmeasured. What is measured is narrower and decides the implementation regardless of how that question resolves.

Direct finding High confidence

Two out of three readers never once looked at the short paragraph at the top of a newsletter. Not skimmed it. Never looked at it.4Nielsen Norman Group67% of users had zero fixations within intros averaging three lines.11 June 2006

Although newsletter introductions averaged only three lines, NN/g measured that 67% of users had zero fixations within them. A prose summary is a block two-thirds of readers will not look at, while still consuming vertical space above the first item and bytes against the size budget.4Nielsen Norman Group67% zero fixations on three-line intros; heatmap concentration on headline openings.11 June 2006 · eyetracking

NN/g: 67% zero fixations on introductions averaging three lines. The finding is specifically an indictment of prose, not of orientation. A list of headline fragments is a different object, and it is precisely the object the same heatmaps show readers fixating on. That distinction is what reconciles the panel's disagreement.4Nielsen Norman GroupThe measured object is a prose intro paragraph, not a scannable index.11 June 2006

Direct finding High confidence

Links that jump you further down the same email do not work on iPhones. Since most people read email on an iPhone, a clickable contents list is mostly decoration.8Mailchimp documentationWarns that contacts "will see the table of contents, but the links won't be clickable".current

Internal anchor links do not act in Apple Mail, Gmail, Outlook or Yahoo on iPhone and iPad, and Litmus places Apple at 62.26% of opens. Anchor clicks also bypass ESP redirect tracking, so a contents block built on them cannot be measured even where it works.8Mailchimp documentationContacts see the contents block; the links are not clickable. On Android they work in Gmail but not Samsung Mail.current9Litmus Email Client Market ShareApple 62.26%, Gmail 27.03%, Outlook desktop 5.83% of opens. >1bn opens.July 2026 · vendor telemetry

Anchor failure spans Apple Mail, Gmail, Outlook and Yahoo on iPhone and iPad; on Android they act in the Gmail app but not Samsung Mail. With Apple at 62.26% of opens, an anchor-based table of contents is a dead control for the majority of the audience. Compounding it, anchor clicks never pass through the ESP redirect, so the variant is untrackable by construction.8Mailchimp documentationVendor documentation of its own feature's failure modes.current9Litmus Email Client Market ShareData July 2026, page current 1 August 2026, >1 billion opens.July 2026

Inference Medium confidence

So: three highlights and a count of what else is inside. Not a paragraph, not a list of everything twice, and no jump links.

Three highlights plus category counts, every link pointing outward to its destination. A full contents list just recreates the flat list a second time, a prose paragraph is the block nobody reads, and anchors are dead on most of the audience. All four members converge on this shape despite disagreeing about why.

Non-exhaustive summary: exactly three highlights plus category counts such as "8 developer tools · 6 agent patterns · 10 updates", with every link outward to the item's destination. Because the satisfy-versus-amplify question is unmeasured, the block should carry distinct tracking parameters so one issue answers it: if index links take meaningful click share while below-index items retain theirs, it amplifies; if index clicks cannibalise, it satisfies.

LimitsNo controlled study measures whether a top summary increases depth of reading or satisfies the reader and reduces it. All four backends searched for one and none found it; the strongest available guidance is Nielsen Norman Group's recommendation of a brief contents block, which is usability guidance rather than a reported effect size. This is the largest single evidence gap bearing on the design.

05; What the inbox actually permits

The rendering rules are not preferences, and one of them deletes your logo

Email is not the web. The rules below are things that break, not things that look worse.

These constraints are the least contested part of the corpus. All four members agree on nearly all of it, and it is where a generating skill can be strictest, because almost every rule here is mechanically checkable.

This section carries the highest inter-member agreement and the highest gate density. Every rule below is verifiable against source or against a render.

Hard rendering constraints and their consequences
ConstraintWhat actually happensSource class
Gmail strips <svg>The tag is removed from the DOM. A vector logo does not degrade, it vanishes.Vendor docs23Google Workspace; Gmail CSS SupportGmail publishes a supported-CSS allowlist and states unsupported properties may be ignored.current
Outlook uses the Word enginemax-width, border-radius and CSS background images ignored; padding unreliable on div and p; margin:0 auto does not centre; media queries ignored; web fonts fall back to Times New Roman rather than to the next font in the stack.Primary + practitioner24Rémi Parmentier; Making sense of Outlook's rendering engineMicrosoft's CSS model splits properties into CORE, COREEXTENDED and FULL, with FULL available only on table elements.2020
Gmail publishes a CSS allowlistNo position, no flex child properties, no grid, no transform, no transition, no animation. var() is supported but the custom-property declaration is not, so every var() resolves to its fallback.Primary vendor docs25Can I Emaildisplay:flex last tested 2 November 2021; CSS custom properties 25 February 2020. Treat both scores as lower bounds with unknown drift.stale test dates
Clipping at ~102KBHTML only, images excluded. Truncates at whatever byte the limit falls on, so tables can be left unclosed. Hides bottom-placed tracking pixels and the unsubscribe footer.Practitioner, undocumented by Google
Threads share a size budgetGmail evaluates same-subject messages as one combined size, so a recurring digest with a stable subject clips as a thread even when each issue is small.Practitioner
Dark mode cannot be opted out ofFull inverters ignore prefers-color-scheme entirely. Outlook.com rewrites markup, injecting data-ogsb and data-ogsc.Vendor guidance
Direct finding High confidence

Gmail cuts long emails off at about 102 kilobytes and hides the rest behind a link. It can take the unsubscribe button with it, which is a legal problem rather than a design one. Google has never published this number anywhere.10hteumeuleu/email-bugs #41Mailchimp, Litmus and Klaviyo all document ~102KB; Google publishes no clipping limit in Gmail Help.ongoing

Gmail clips at approximately 102KB of HTML, and Google documents no such limit anywhere. Mailchimp, Litmus and Klaviyo all describe it independently. Clipping truncates mid-markup, so tables may be left unclosed, and it hides bottom-placed tracking pixels and the unsubscribe footer.10hteumeuleu/email-bugs #41, Gmail clips emails at 102 kBPractitioner-verified across three vendors. Google's only published Gmail size limit is the 25MB attachment cap.ongoing

The limit applies to the HTML portion only; images, attachments and fonts do not count against it. Three consequences specific to a multi-item digest: truncation lands at an arbitrary byte and may leave tables or divs unclosed; bottom-placed tracking pixels including the ESP open pixel may never load; and the visible unsubscribe link can disappear, which is a compliance failure rather than a cosmetic one.10hteumeuleu/email-bugs #41Google documents no clipping limit anywhere in Gmail Help.ongoing11Litmus, How to Keep Gmail from Clipping Your EmailsDocuments the threading behaviour: same-subject messages evaluated as a combined size.ongoing

LimitsThe threshold is practitioner-observed and corroborated across three vendors rather than published, so it could move without notice. Klaviyo separately documents mobile clients clipping far earlier, around 20KB on iOS and 75KB elsewhere. No other vendor corroborates that, and if it were true a 24-item HTML digest would be impossible at any reasonable markup weight, so it is recorded as unverified rather than designed around.

Direct finding High confidence

The dark-mode meta tags are a promise, not a hint. Apple Mail leaves your colours alone if the tags are absent, but partially inverts them if the tags are present without matching dark styles. Declaring support you have not built is worse than declaring nothing.12Litmus, Ultimate Guide to Dark Mode for Email MarketersApple Mail partially inverts when the colour-scheme meta tags are present without accompanying dark styles.ongoing

color-scheme and supported-color-schemes must co-occur with a prefers-color-scheme: dark block, or be omitted entirely: Apple Mail leaves markup untouched without the tags and partially inverts with the tags and no dark styles. This is a co-presence assertion a linter can make, and one of the few dark-mode rules that is deterministic rather than heuristic. The rest is defensive: Outlook.com's partial inversion targets #000000 and #FFFFFF specifically rather than reacting to lightness, so near-black and near-white sidestep it.12Litmus, Ultimate Guide to Dark ModeFull inverters include Gmail iOS, Outlook 2021 for Windows, Office 365 for Windows and Windows Mail.ongoing

LimitsThe near-value workaround is a heuristic against an undocumented detection rule, not a guarantee. Dark-mode adoption figures also do not reconcile: Litmus variously reports over 40%, more than 35%, and a 35% average for 2022, and every one of those is an Apple-only measurement being reported as a whole-audience share. No clean current figure was located by any member.

06; The picture rule everyone repeats

The sixty-forty ratio is not a deliverability rule, and one backend wanted it gated

There is a famous rule that an email needs at least sixty per cent text and no more than forty per cent pictures, or spam filters will catch it. Somebody tested it. It is not true.

One member proposed enforcing a 60:40 text-to-image ratio as a hard gate. The other three rejected it, one with a controlled test. This is the clearest case in the panel of a widely-repeated rule with nothing behind it.

The ratio rule was proposed as a gate by one member and rejected by three, one of them with a controlled test against a filter population.

Direct finding High confidence

Email on Acid ran emails past twenty-three spam filters. Above five hundred characters, the balance of text to pictures made no difference to whether they got through.13Email on Acid, Does Text to Image Ratio Affect Deliverability?Controlled test against 23 spam filters. Every 500+ character email passed every filter regardless of image count.undated

Email on Acid tested against 23 popular spam filters and found that at 500 characters or more, content-to-image ratio does not affect deliverability. Every 500-plus-character email passed every filter regardless of image count. The quoted rule also varies between 60/40 and 80/20 depending on the source, which is itself diagnostic that no filter enforces a number.13Email on AcidControlled test, 23 spam filters.undated14Badsender, Text to image ratio in an emailFiled under "Myths and Legends of Deliverability", attributing its persistence to listicle recycling.8 October 2020

Email on Acid, controlled test against 23 popular spam filters: at 500+ characters, content-to-image ratio does not affect deliverability. Badsender files the rule under "Myths and Legends of Deliverability" and attributes its persistence to listicle recycling. The variance in the quoted figure between 60/40 and 80/20 is itself evidence that no filter enforces one.13Email on AcidEvery 500+ character email passed all 23 filters regardless of image count.undated

LimitsApache SpamAssassin does contain image-only and low-text-to-image heuristic tests, so extreme image dependence can contribute to that scorer. SpamAssassin is one open-source scoring system rather than a description of how Gmail, Microsoft or Yahoo classify, and the residual real risk is indirect: an image-only email that renders blank prompts complaints, and complaints are a reputation signal.

Inference High confidence

Gate the real thing instead. A picture may carry mood. It may never carry meaning, because for some readers it will simply not be there: Outlook blocks images by default and will not even let the fallback text be styled.15Campaign Monitor; Displaying and Optimizing ALT TextOutlook desktop shows alt text only after a security warning and will not allow it to be styled at all.undated

Gate against image-only communication rather than against a ratio. Outlook desktop blocks images by default, will not let alt text be styled at all, and shows it only behind a security warning. Several clients reject alt text that exceeds the image width. Every one of those failure modes lands on the same element, so a featured item's headline must exist as HTML text beside the banner, never inside it.15Campaign Monitor, Displaying and Optimizing ALT TextOutlook desktop 2007-2016 shows alt text only after a security warning and will not allow it to be styled at all.undated

The banner's failure modes are blocked, broken, width-clipped alt and unstylable alt, and they destroy the headline and the AI-generated inbox summary simultaneously. Alt text also renders when images are turned off but not when an image is genuinely broken, so alt is not a fallback for a dead URL. The rule is a text-bearing heading as a sibling to the <img>, not a ratio.15Campaign MonitorGmail underlines linked-image alt text and overrides its colour with blue; Yahoo underlines it.undated

LimitsImage-blocking prevalence is structurally unmeasurable and always will be: open detection depends on an image loading, so a reader with images blocked is invisible to the instrument. The widely-circulated "43% of users view email with images off" is a 2013 Litmus figure specific to Gmail before it switched images on by default, and Litmus itself cautions against extrapolating from it.

Direct finding Medium confidence

The middle tier is the weak spot. NN/g found readers preferred full-width imagery and rated thumbnails "less valuable and compelling", and re-classified a newsletter it had previously praised as "cluttered" on re-test. So the compact rows should be text-forward with decorative icons, and must read correctly with every image stripped.16Kim Flaherty, NN/g, Marketing Email and Newsletters: UX Findings Then and NowDiary study n=9 plus usability testing n=28. Thumbnails rated less valuable than full-width photos.13 August 2017

NN/g's re-test found full-width, high-quality photos preferred and thumbnail imagery rated less valuable and compelling, with a previously-praised thumbnail newsletter re-classified as cluttered. The compact tier's small icons are thumbnails by another name. The resolution is asymmetric rather than a choice between evidence bases: large imagery only in the featured tier, where the finding applies and the item count is small, with the compact tier in the developer-newsletter idiom of text rows carrying alt="" decorative icons.16Kim Flaherty, NN/gDiary study n=9 + usability testing n=28.13 August 2017

LimitsGeneral consumer audiences, not developers. It cuts directly against the developer-newsletter convention, which is close to imageless and is demonstrated at scale by three publishers at 180,000, 216,000 and 1.6 million subscribers, and is corroborated by the one first-party observation from a vendor whose own audience is software engineers.29Email on Acid; Plain Text vs HTML EmailsReports from its own sending to software engineers and IT professionals that readers respond better to less pushy, more plain-text-like design.undated · first-party These are not reconcilable by picking one, because the NN/g finding is measured and the developer convention is unmeasured but demonstrated. No published test of banner imagery against text-only for a developer audience exists.

07; The part that is already solved

Nearly every email fails accessibility, and nearly every failure is mechanical

Someone checked 443,585 real emails against automated accessibility tests. Twenty-one passed.

The accessibility picture is bleak and, for once, that is good news: the failures are dominated by four mechanically detectable defects, which means they are all gateable.

Near-universal failure across a large corpus, concentrated in mechanically detectable defect classes. This is a tooling gap rather than a hard problem.

Direct finding High confidence

Of 443,585 emails checked, 99.89% had serious or critical problems. Twenty-one were clean.17Email Markup Consortium, Accessibility Report 2025443,585 HTML emails collected May 2024 to May 2025. Only 21 passed every automated check.13 October 2025

The Email Markup Consortium analysed 443,585 HTML emails collected between May 2024 and May 2025 and found 99.89% contained automated issues rated serious or critical. Missing roles on layout tables in 86.24%, missing alt text in 51.42%, links without discernible text in 72.04%, insufficient contrast in 59.37%.17Email Markup Consortium, Accessibility Report 2025n=443,585, automated checks, published methodology.13 October 2025

EMC Accessibility Report 2025, n=443,585 collected May 2024 to May 2025, automated checks only: 99.89% serious or critical, 21 emails passing cleanly. The failure distribution matters more than the headline, because the four most common defects (layout-table roles 86.24%, alt text 51.42%, discernible link text 72.04%, contrast 59.37%) are all mechanically detectable and therefore all gateable.17Email Markup ConsortiumAutomated checking only, which catches mechanical failures rather than the full standard.13 October 2025

LimitsAutomated checking only, which catches mechanical failures rather than the full standard. A page passing every automated check is not thereby accessible.

Inference High confidence

Screen-reader users move through a long document by jumping between its headings. If your tiers are only a visual effect, those users get nothing from them.18WebAIM Screen Reader User Survey #1071.6% navigate long pages by headings, rising to 78% among advanced users.December 2023 - January 2024

WebAIM found 71.6% of screen-reader users navigate long pages by headings, rising to 78% among advanced users. Both audiences are doing the same thing: traversing by landmark and rejecting in bulk. A tiered design built from styled divs delivers the benefit to sighted readers and nothing at all to screen-reader users, which is the accessibility failure most likely to occur here and the easiest to lint for.18WebAIM Screen Reader User Survey #1071.6% navigate by headings; 88.8% rate heading levels very or somewhat useful.December 2023 - January 2024

Heading navigation is the screen-reader equivalent of visual tiering, so the tier boundaries must be the heading outline: one <h1>, an <h2> per tier, no skipped levels, with margin:0 and mso-line-height-rule:exactly inline to defeat Word's default heading margins. role="presentation" is not inherited by nested tables, so every layout table needs it individually.18WebAIM Screen Reader User Survey #10Large-sample survey; 88.8% rate heading levels useful.December 2023 - January 2024

LimitsContrast is the one accessibility rule a client can undo unilaterally: a palette passing WCAG's 4.5:1 in light mode has no guaranteed ratio after a full inverter repaints it, and no dataset quantifies how often that produces a failure.28W3C; WCAG 2.2 Success Criterion 1.4.34.5:1 for normal text, 3:1 for large text.WCAG 2.2 · normative WebAIM's data is also about web pages rather than email. The mechanism transfers, since a digest is precisely a long document with repeated structure, but the measurement was not taken in an inbox.

Inference Medium confidence

For a recurring digest, "is it actually read" stops being a design goal and becomes a deliverability constraint. Not because low clicks are punished directly, which nobody documents, but because an unreadable recurring send accumulates complaints, and Google publishes a complaint threshold with consequences attached.19Google Workspace Admin Help, Email sender guidelinesSpam rates must stay below 0.30%, with guidance to stay below 0.10%. Effective 1 February 2024.current

Two steps, only one of which is documented. Step one: Google requires spam rates below 0.30% with guidance to stay under 0.10%, measured in Postmaster Tools, recalculated daily with graduated consequences rather than a cliff. Step two, undocumented by any mailbox provider: that low positive engagement reduces reputation independently of complaints. Stated as a two-step inference rather than a measured relationship, because the second step has no primary source.19Google, Email sender guidelinesFor senders above 5,000 messages/day: SPF, DKIM, DMARC, alignment, RFC 8058 one-click unsubscribe, and a visible body unsubscribe link.Effective 1 February 202420RFC 8058Signaling One-Click Functionality for List Email Headers.January 2017

LimitsMailbox providers do not publish a causal rule saying that a low click rate by itself reduces sender reputation. The documented pathways are complaints, authentication failures, bounces and list practice. The chain here is two steps and only the first is sourced.

08; What nobody can source

Eight numbers everyone repeats and nobody can trace

A lot of email advice rests on statistics that, when you go looking, have no study behind them.

Two members independently ran traceability checks on circulating figures. Eight came back unsourceable. They are listed because the generating skill should refuse to quote them, and because the pattern behind them is instructive.

Traceability failures, recorded so they can be refused rather than rediscovered.

The most instructive one is the fold. The article everyone cites for email fold statistics takes every number from research about web pages, carried out between 2010 and 2018. Litmus flags its own post as aged, and its actual conclusion is "it depends". Any email fold statistic, traced back far enough, arrives at web data.21Litmus, The fold debateSources all figures from NN/g web page usability research 2010-2018. No email-specific scroll-depth measurement. Litmus flags the post as 2+ years old.14 May 2021

The fold case is the instructive one. Litmus's fold article is the most-cited source in the field and takes all of its figures from NN/g web page usability research covering 2010–2018: that 43% of time is spent below the fold and 81% falls in the first three screenfuls. There is no email-specific scroll-depth measurement in it, Litmus flags the post as aged, and its conclusion is "it depends" with the fold reframed as "a few swipes of the thumb". Scroll depth is not natively trackable in email at all, which is why the proxy that publishers actually report is click distribution by template position.21Litmus, The fold debate: Do emails still need to stay above the fold?Vendor analysis, flagged aged by the publisher itself.14 May 2021

Inference Medium confidence

Two things follow for the subject line. Say what is actually inside, because a bare number tells the reader how much work you are about to give them and nothing about whether any of it matters. And the little preview line under the subject is increasingly written by Apple's software rather than by you.27Litmus, AI-Generated Email SummariesApple Intelligence pre-open summaries are on by default and replace preview text with a two-line summary built from the message HTML.2024

Name a specific capability and put the count second: the only causal evidence shows that naming a relevant item raises that item's detail-reading, while a bare count sets a scope expectation without a relevance one. Treat subject length as a truncation constraint and warn rather than fail on it, because three large datasets give three different optima and an academic study of 455 million users found no direct relation at all.26Review of Managerial Science, May I have your attention, please?5,765 promotional emails, 455m users, 73 countries. Found no direct relation between subject length and attention.2022 · peer-reviewed Meanwhile the preheader has stopped being a controlled surface: Apple Intelligence generates the inbox preview from your HTML, and Apple is 62.26% of opens.27Litmus, AI-Generated Email SummariesOn by default. Summary quality is reported as mixed, and worst for image-heavy mail.2024

Kong et al. found that featuring a relevant item in the subject raised that item's read-in-detail rate from 15% to 24% while producing no significant overall open-rate difference. Length guidance is unusable as a lever: GetResponse puts opens highest at 61–70 characters and CTOR highest at 41–50, Attentive puts the optimum at 25–35, and Adestra found CTOR optimised above 70. None controls for sender, industry or offer.26Review of Managerial SciencePeer-reviewed null: no direct relation between subject length and attention across 455m users.2022 Separately, with Apple at 62.26% of opens and Apple Intelligence generating the inbox preview from the message HTML, the preheader field is best treated as a fallback and the first real body text as the actual input, which compounds the live-text requirement.27Litmus, AI-Generated Email SummariesVendor analysis of a primary Apple feature.2024

LimitsNo study isolates numbered against curiosity or benefit framing for a developer audience with a click-side metric. Every claim circulating on that question is an agency-blog assertion with no dataset behind it. There is also no measurement of how often Apple's generated summary displaces an authored preheader in practice.

Direct finding High confidence

The same problem sits under the whole metrics layer. Apple's Mail Privacy Protection downloads remote content whether or not a human looked, and Apple is 62.26% of opens. That breaks open rate, and it takes click-to-open down with it, because click-to-open divides by the same broken number.22Apple, Mail Privacy ProtectionProtected Mail clients download remote content in the background regardless of engagement.current · primary vendor documentation9Litmus Email Client Market ShareApple at 62.26% of opens; MPP opens "are not considered reliable opens"; MPP estimated to touch 55-60% of all opens.July 2026

Apple documents that protected Mail clients download remote content in the background regardless of engagement. Litmus places Apple at 62.26% of opens and states MPP opens are not considered reliable, estimating MPP touches roughly 55–60% of all opens. Pixel-derived read time is likewise unavailable for that audience: Litmus writes -1 into its read_seconds field. Native scroll depth does not exist in email, because HTML mail cannot execute the instrumentation the web uses. The usable metric is bot-filtered unique clickers per delivered email.22Apple, Mail Privacy ProtectionPrimary vendor privacy documentation.current

LimitsClicks are better but not pristine. Security and privacy scanners click links to check for malware, which can inflate raw click rates by a reported 10–35% on B2B lists, so a unique click must mean a bot-filtered click. And a bounded audit by one member found six of six prominent subject-line guidance pages still optimising for open rate, so the contaminated metric remains in active use across the industry's own benchmarks.

09; The panel

Four backends, one brief, and a disagreement worth the money

Five research systems were commissioned to answer the same question independently. Four finished. Where they disagreed is where the useful answers were.

The panel exists so that agreement means something. It earned its cost twice over here, because the members split on the central question and the split forced a source-check that a single opinion would never have prompted.

Panel composition, completion state and cost. Support is counted in independent registrable domains rather than in how many members agreed.

Four material disagreements, each resolved against the better-sourced member rather than by majority: whether an item ceiling exists, whether a contents block suppresses engagement, whether the text-to-image ratio should be gated, and where the click-decay figures came from. In three of the four, the member holding the minority position was the one with primary sources.

Methods, and what this page could not establish

Panel
Five backends commissioned on one brief through Dossier. Four completed: OpenAI gpt-5.6, Claude Code, Gemini (fast tier) and Perplexity Sonar Deep Research. A Gemini max-tier run was abandoned after 165 minutes; a cancel was issued, took effect locally, and reverted to running sixty seconds later because the provider still held the task, so the run was left alive and excluded rather than stopped. An Antigravity CLI lane never started because its binary could not be identified. 182 cited sources across the four completed reports.
What was read
Every completed report was exported and read end to end rather than from the merged distillation, which is a coverage difference between reports rather than a summary of them. The registry below lists the sources this page cites directly, which is a subset of the corpus.
What was verified
Where members disagreed about a figure, the disagreement was resolved by tracing the figure to its origin. That is how the click-decay numbers were found to be search-engine data, and how the fold statistics were found to be web-page data. Eight further figures were checked and could not be traced at all; they are listed in section 08.
What a human reviewed
The brief, the decision to commission a paid panel, and the finding. The synthesis, claim graph and page are machine-authored and machine-designed, then read and published by a person under their own name.
Deliberate substitutions
The Mobbin MCP reference pass and the /trawl aesthetic divergence pass prescribed by the publishing skill were not run; the layout derives from the claim graph and the house reading-surface rules instead. No figure on this page is generated imagery. The argument has no scrubbed or pinned episode, so the motion layer is spent on entrance choreography and interface feedback, which the skill permits provided it is recorded here.
What this page could not establish
Four of the seven questions in the brief have no published study behind them: whether a tiered layout beats a flat list of the same items; whether a top summary block increases depth of reading or reduces it; whether banner imagery helps or hurts a developer audience; and what an email-native scroll-depth benchmark would show. Three of those are instrumentation-blocked rather than merely unstudied. Image-blocking prevalence cannot be measured, because open detection depends on an image loading. Scroll depth has no native event model in email. Anchor-link engagement is untrackable because anchors bypass the ESP redirect. These will not be resolved by better searching.
Cost
$23.00 committed across five reservations against a $250 daily ceiling. The abandoned Gemini max run accounts for $7.00 of that and produced nothing. A restart attempt de-duplicated onto the still-running original and was not charged.

Registry

Every source this page cites, and what kind of evidence it is

The four reports cite 182 sources between them. Listed below are the sources this page draws on directly, each with the class of evidence it represents, because a peer-reviewed field experiment and a vendor blog post are not interchangeable even when they agree.

  1. MailerLite; How Many Links Should An Email HaveObservational · 317,000 campaigns / 2.9bn emails · 12 February 2026
  2. Scheibehenne, Greifeneder & Todd; Journal of Consumer Research 37(3) 409–425Meta-analysis · 63 conditions / 50 experiments / N=5,036 · 2010
  3. Chernev, Böckenholt & Goodman; Journal of Consumer PsychologyCompeting meta-analysis · 99 experiments / 53 papers · 2015
  4. Nielsen Norman Group; Email Newsletters: Surviving Inbox CongestionEyetracking study · n=42, 117 newsletters · 11 June 2006
  5. Kumar & Salo; Journal of Marketing Communications 24(5) 535–548Peer-reviewed · LianaMailer analytics · 16 March 2016
  6. Kong et al.; ACM CSCWPeer-reviewed field experiment · 8 weeks, 117 participants, 4,242 messages · 11 November 2022
  7. Mailchimp; Add a Table of Contents to Your EmailVendor documentation of its own feature's failure modes · current
  8. Litmus; Email Client Market ShareVendor telemetry · >1 billion opens · data July 2026
  9. hteumeuleu/email-bugs #41; Gmail clips emails at 102 kBPractitioner-verified · undocumented by Google · ongoing
  10. Litmus; How to Keep Gmail from Clipping Your EmailsVendor rendering guidance · ongoing
  11. Litmus; Ultimate Guide to Dark Mode for Email MarketersVendor rendering guidance · ongoing
  12. Email on Acid; Does Text to Image Ratio Affect Deliverability?Controlled test against 23 spam filters · undated
  13. Badsender; Text to image ratio in an emailPractitioner analysis · 8 October 2020
  14. Campaign Monitor; Displaying and Optimizing ALT Text in Popular Email ClientsClient rendering test data · undated
  15. Kim Flaherty, NN/g; Marketing Email and Newsletters: UX Findings Then and NowDiary study n=9 + usability testing n=28 · 13 August 2017
  16. Email Markup Consortium; Accessibility Report 2025Automated audit · n=443,585 emails · 13 October 2025
  17. WebAIM; Screen Reader User Survey #10Large-sample survey · December 2023 – January 2024
  18. Google Workspace Admin Help; Email sender guidelinesPrimary vendor requirements · effective 1 February 2024
  19. RFC 8058; Signaling One-Click Functionality for List Email HeadersPrimary technical standard · January 2017
  20. Litmus; The fold debate: Do emails still need to stay above the fold?Vendor analysis · flagged aged by the publisher · 14 May 2021
  21. Apple; Mail Privacy ProtectionPrimary vendor documentation · current
  22. Google Workspace; Gmail CSS SupportPrimary vendor documentation · current
  23. Rémi Parmentier; Making sense of Outlook's rendering enginePractitioner analysis of Microsoft's CSS model · 2020
  24. Can I Email; display:flex, and CSS VariablesClient rendering test data · last tested 2 November 2021 and 25 February 2020 respectively
  25. Review of Managerial Science; May I have your attention, please?Peer-reviewed · 5,765 emails, 455m users, 73 countries · 2022
  26. Litmus; AI-Generated Email Summaries: What Marketers Need To KnowVendor analysis of a primary feature · 2024
  27. W3C; WCAG 2.2 Success Criterion 1.4.3 Contrast (Minimum)Normative accessibility standard
  28. Email on Acid; Plain Text vs HTML EmailsFirst-party practitioner observation, technical audience · undated

Research that argues with itself is worth more than research that agrees.

This page exists because four independent research backends were given one brief and came back disagreeing about the central question. Three recommended cutting the item count. The two best-sourced showed there is no evidence for an item ceiling at all. Dossier runs the panel, keeps every member's report whole, and makes the disagreement the finding rather than averaging it away. Margin is where the argument gets written down.

Panel
5 backends commissioned, 4 completed, 1 abandoned, 1 never started
Corpus
182 cited sources across 4 independent backends
Spend
$23.00 committed against a $250 daily ceiling
Disagreements
4 material, all resolved against the better-sourced member
Unmeasured
4 of the 7 questions have no published study behind them