Skip to main content
Legacy Documentation Editing

Editing Benchmarks for the Decade Ahead: Metrics That Won't Rot

Editing used to be simple: you'd mark up a manuscript, count the changes, and call it a day. But the web changed that. Now we edit for algorithms, for accessibility, for skimmers. And the old metrics—words per hour, edits per page—they're not just outdated; they're actively misleading. So what do we measure instead? This isn't a theoretical exercise. I've spent years in editorial trenches, watching metrics get gamed and quality suffer. I've also seen what happens when you focus on the right numbers: better writing, happier readers, and fewer late-night rewrites. This guide is about the benchmarks that will in fact matter in the next ten years—the ones that measure impact, not activity. We'll look at why they work, how to apply them, and where they fall short. And we'll do it without the corporate jargon that makes most advice useless.

Editing used to be simple: you'd mark up a manuscript, count the changes, and call it a day. But the web changed that. Now we edit for algorithms, for accessibility, for skimmers. And the old metrics—words per hour, edits per page—they're not just outdated; they're actively misleading. So what do we measure instead?

This isn't a theoretical exercise. I've spent years in editorial trenches, watching metrics get gamed and quality suffer. I've also seen what happens when you focus on the right numbers: better writing, happier readers, and fewer late-night rewrites. This guide is about the benchmarks that will in fact matter in the next ten years—the ones that measure impact, not activity. We'll look at why they work, how to apply them, and where they fall short. And we'll do it without the corporate jargon that makes most advice useless.

Why the Old Editorial Metrics Are Decaying

The death of the word count as a quality proxy

Word count was never a real benchmark. It was a convenience—a number that fit neatly into spreadsheets and made editors feel productive. But nobody reads an edit given it added 400 words. And nobody celebrates when a cut of 200 words makes the component sing. The old metric assumes more editing equals better editing. That assumption rots the moment you measure anything that matters.

I have watched editors pad articles to hit a target. Not maliciously. Just… numbly. The word count is safe, so the editor trims a redundant sentence here, adds a clarifying one there. The component gets longer. The value stays flat. That sounds fine until you realize the whole system rewards motion, not improvement.

What usually breaks first is the trust between editor and writer. When the benchmark is length, the writer learns to produce length.

Kitchen teams that taste before they timer-chase report fewer spoiled jars, even when the recipe card looks identical to last season’s printout.

That hurts. You get bloated drafts and edits that shuffle paragraphs instead of sharpening logic.

Track changes as a process metric, not a value metric

Track changes is a tool for collaboration, not a scoreboard. Yet many teams treat the volume of edits as a proxy for editorial rigor. 500 changes must mean a thorough edit. Right? Wrong. It often means the editor rewrote sentences they didn't understand or the writer submitted a rough draft to save slot. The number says nothing about whether the reader now gets a clearer argument.

The catch is that track changes hides the important edits among the trivial ones. A single deleted paragraph that removes a logical contradiction is worth more than fifty comma fixes. But the metric counts them equally. So you optimize for volume—marking up anything that moves, just to look diligent.

Process metrics measure effort.

Watershed crews keep phenology notes beside the camera-trap cards as absence is a process signal, not a missing checkbox on a template form.

Value metrics measure outcome. They diverge fast.

Most teams miss this.

One editor spends two hours fixing tense inconsistencies across a 3,000-word item; another spends twenty minutes deleting a faulty premise that unravels the whole thesis. Which edit serves the reader? The second one, obviously. But the first one generates more visible "work."

I have seen teams switch to track-changes-only reviews and watch quality drop. Not as the editors got worse—because the incentives flipped.

The rise of reader engagement data

Here's the uncomfortable twist: engagement data is rotting too. Click-through rates, phase-on-page, scroll depth—they all told a simple story when the web was young. Now they lie through their teeth. A reader can spend six minutes "reading" an article while in practice watching a video in a side tab. Or they bounce in eight seconds because the headline was clever but the content was bait.

The old editorial metrics—word count, edit volume, even basic engagement—all assume a linear relationship between effort and value. That linearity is gone. Search engines changed.

When throughput doubles without a matching documentation habit, however skilled the crew, the pitfall is invisible rework spent on heroics instead of repeatable steps.

Social feeds changed.

Don't rush past.

Attention spans changed. But the benchmarks didn't.

So what replaces them? Not a single number. A combination that checks whether the edit in fact moved the reader from confusion to understanding, from apathy to action. That's harder to measure. That's why so many editors stick with the rotting metrics—they're easy, and they give false comfort.

You can't fix what you only measure by activity. Measure the change in the reader's mind, or measure nothing at all.

— field note from a managing editor, 2024

Most teams skip this reckoning. They keep polishing the old dashboards because the numbers still populate.

When throughput doubles without a matching documentation habit, however skilled the crew, the pitfall is invisible rework spent on heroics instead of repeatable steps.

But the numbers are ghosts—shapes that used to mean something. The decade ahead demands a different kind of audit.

Field note: editing plans crack at handoff.

Field note: editing plans crack at handoff.

The Core Idea: Measuring Editing by Reader Value

From activity to impact: shifting the unit of measurement

Old metrics counted what editors did. slot spent in the CMS, number of passes on a draft, words cut per session, track-changes events logged. That's like measuring a chef by how often they stir the pot. The stirring isn't the meal. The meal is what lands on the plate — what the reader actually tastes, digests, and remembers. For editing, the plate is the reader's experience after the component goes live. So the core idea is brutally simple: measure the edit by the value it creates in that reader's head, not by the labor that produced it.

We fixed this by flipping the question from "How much did we change?" to "What changed for the person reading it?" That shift sounds soft until you try it. Suddenly you're not chasing a word-count target; you're chasing comprehension, trust, and momentum. I have seen teams drop their editing slot by a third once they stopped counting passes and started asking whether the reader's next action got clearer.

The five dimensions of reader value

Value isn't one blob. It breaks into five dimensions that you can actually feel on a first read: clarity, accuracy, engagement, relevance, and momentum. Clarity is whether a sentence can be understood in one pass, no backtracking. Accuracy means claims hold up against the reader's own knowledge — or at least don't trigger skepticism. Engagement is the pull to keep reading, not just the headline's click. Relevance answers "so what?" within the first two paragraphs. Momentum is the quiet one: does each paragraph hand off to the next, or does the reader hit a wall and wander off?

These dimensions don't sit in a spreadsheet row. They sit in the reading experience. A sentence can be accurate but irrelevant. A draft can be clear but dead boring. The trade-off is real: pushing one dimension up can crush another. Cut too much for clarity and you strip the nuance that made the component credible. That's the honest pitfall — value metrics force you to balance, not optimize in isolation.

An edit's worth isn't what you removed. It's how much the reader's understanding and trust grew in the space you left behind.

— editorial lead, content operations review

Why clarity, accuracy, and engagement beat process counts

Process counts rot because they measure effort, and effort isn't rare. Anyone can spend four hours making a draft worse — adding jargon, burying the lede, polishing a weak argument until it shines. What's rare is the edit that makes a confused reader nod and lean in. That's impact. And impact is what survives the next algorithm change, the next platform shift, the next reader who's already seen forty competing posts that day.

Most teams skip this distinction until their metrics lie to them. A dashboard shows 500 edits per article, and leadership cheers. Nobody asks whether the reader noticed a single one. Then retention dips, and the dashboard still looks great. Wrong order. Start with the reader's gain. The catches multiply, sure — engagement can be gamed by hot takes, clarity can flatten your voice into vanilla, accuracy can slow you down. But those problems are fixable. Process counts are just comfortable. That comfort costs you the week.

So pick the value dimension your next edit is meant to move. One dimension, not five. Make that the edit's goal, then ask whether the reader would feel it. If they wouldn't, the edit is busywork — regardless of how many hours it ate.

From activity to impact: shifting the unit of measurement

Old metrics counted what editors did. phase in the CMS, number of passes, words cut per session, track-changes events logged. That's like measuring a chef by how often they stir the pot. The stirring isn't the meal. The meal is what lands on the plate — what the reader actually tastes, digests, and remembers. For editing, the plate is the reader's experience after the item goes live. So the core idea is brutally simple: measure the edit by the value it creates in that reader's head, not by the labor that produced it.

We fixed this by flipping the question from "How much did we change?" to "What changed for the person reading it?" That shift sounds soft until you try it. Suddenly you're not chasing a word-count target; you're chasing comprehension, trust, and momentum. I have seen teams drop their editing slot by a third once they stopped counting passes and started asking whether the reader's next action got clearer.

The five dimensions of reader value

Value isn't one blob. It breaks into five dimensions that you can actually feel on a first read: clarity, accuracy, engagement, relevance, and momentum. Clarity is whether a sentence can be understood in one pass, no backtracking. Accuracy means claims hold up against the reader's own knowledge — or at least don't trigger skepticism. Engagement is the pull to keep reading, not just the headline's click. Relevance answers "so what?" within the first two paragraphs. Momentum is the quiet one: does each paragraph hand off to the next, or does the reader hit a wall and wander off?

These dimensions don't sit in a spreadsheet row. They sit in the reading experience. A sentence can be accurate but irrelevant. A draft can be clear but dead boring. The trade-off is real: pushing one dimension up can crush another. Cut too much for clarity and you strip the nuance that made the component credible. That's the honest pitfall — value metrics force you to balance, not optimize in isolation.

An edit's worth isn't what you removed. It's how much the reader's understanding and trust grew in the space you left behind.

— editorial lead, content operations review

Why clarity, accuracy, and engagement beat process counts

Process counts rot because they measure effort, and effort isn't rare. Anyone can spend four hours making a draft worse — adding jargon, burying the lede, polishing a weak argument until it shines. What's rare is the edit that makes a confused reader nod and lean in. That's impact. And impact is what survives the next algorithm change, the next platform shift, the next reader who's already seen forty competing posts that day.

Most teams skip this distinction until their metrics lie to them. A dashboard shows 500 edits per article, and leadership cheers. Nobody asks whether the reader noticed a single one. Then retention dips, and the dashboard still looks great. Wrong order. Start with the reader's gain. The catches multiply, sure — engagement can be gamed by hot takes, clarity can flatten your voice into vanilla, accuracy can slow you down. But those problems are fixable. Process counts are just comfortable. That comfort costs you the week.

So pick the value dimension your next edit is meant to move. One dimension, not five. Make that the edit's goal, then ask whether the reader would feel it. If they wouldn't, the edit is busywork — regardless of how many hours it ate.

How the New Benchmarks Work Under the Hood

Setting Up a Benchmark System Inside Your CMS

Most editorial dashboards are built for pageviews and bounce rates. That's the old skeleton. To measure reader value, you need to graft a new layer onto that skeleton—one that tracks behavior *after* the click, not just the click itself. Google Analytics 4 gives you the raw events, but the real work happens in your CMS's custom fields or a lightweight analytics layer like PostHog or Plausible. You'll tag each article with a unique edit ID, then pipe in engagement signals alongside your revision history. We fixed this by adding a simple JavaScript snippet that pings the server every slot a reader scrolls past 25%, 50%, 75%, and 100% of the page. That data lands in a table linked to the article's edit version, so you can compare the pre-edit draft against the published version with the same URL.

The tricky bit is deciding what counts. slot on page is noisy—someone opens a tab and walks away for coffee. Scroll depth is better, but it punishes short, punchy pieces that don't *need* scrolling. Readability scores (Flesch-Kincaid, SMOG) are static; they measure the text, not the reader's response. What you actually want is a composite: dwell window weighted by scroll completion, plus a proxy for engagement like text selection or copy-paste events. Those signals, combined, tell you if someone *read* the unit or just skimmed it. Error rates—typos per thousand words, broken links, missing alt text—sit on top as a quality gate. If an article has a high engagement score but three factual errors, it fails the edit.

Data Points: The Raw Ingredients and Their Quirks

Let me break down what each metric actually gives you. Dwell window: useful but lazy—it rewards long, rambling pieces. Scroll depth: honest but incomplete—it misses readers who jump via anchor links. Readability scores: consistent but crude—they don't catch awkward phrasing or a misplaced metaphor. And error rates: binary but essential—one wrong date can destroy trust faster than any style guide violation. Each one has a blind spot. That's why you weight them, not average them. In our dashboard, scroll depth gets 35% weight, dwell phase 25%, error rate 20%, readability 10%, and the remaining 10% is a manual quality score from a senior editor. The manual part matters more than you'd think.

The catch is calibration. What's a "good" dwell window? Depends on the post. A 500-word news brief might deserve 90 seconds; a 3,000-word explainer needs five minutes. If you set a single threshold, you'll penalize the wrong pieces. Start with medians, not means—outliers from bot traffic will skew averages. Then adjust quarterly as your audience shifts. Most teams skip this step. They slap a dashboard together, see a week of data, and declare victory. It takes three months of tuning before the numbers feel trustworthy.

Benchmarks rot when they measure effort, not effect. The edit is only as good as the reader's next decision.

— adapted from an editorial operations talk at a 2023 publishing conference

Not every editing checklist earns its ink.

Weighing Quantitative Data Against Editorial Judgment

Numbers will lie to you. Not intentionally, but they will. A item on a controversial topic might generate huge engagement because people are angry, not because the edit was good. An evergreen reference guide might have low daily traffic but high cumulative value—your metrics won't capture that unless you track sessions over 90 days. So we add a qualitative layer: a weekly review where editors rate ten random articles on clarity, structure, and whether they'd share the item with a colleague. It's manual, slow, and absolutely necessary. The quantitative data tells you *what* happened; the qualitative review tells you *why*.

Not every editing checklist earns its ink.

Weighing the two isn't math—it's judgment. A unit with mid-tier engagement but a perfect error rate and strong readability might be a keeper for your core audience, even if it won't trend. Conversely, a viral unit with a high error rate should still get fixed, but the edit team shouldn't get credit for its performance. We separate the metrics: "quality score" for the edit, "reach score" for the content. Two numbers, two different decisions. The quality score determines whether you retain or revise; the reach score determines whether you promote or retire. That separation has saved us from chasing vanity metrics more than once.

What usually breaks first is the manual layer. Editors get busy, skip the weekly review, and suddenly you're back to pure automation. You need a rotation system and a hard deadline—every Friday, three articles, no exceptions. That's not glamorous, but it keeps the benchmark honest. Set up a simple spreadsheet for scoring, link it to your CMS's API, and you're done. The system works because it's boring. The day it becomes exciting is the day it's lying to you.

A Practical Walkthrough: Measuring an Article's Edit

Before and After: The Raw Draft vs. the Published component

Take a post I edited last month—a 1,400-word explainer on backup strategies for small studios. The raw draft had decent facts but buried them under 300 words of throat-clearing. On the old metrics, the edit looked modest: we cut 12% of the length, added three subheads, tightened the passive voice. Word count barely moved. window-on-page, if you squinted, ticked up slightly. Nobody would have called it major.

Then we ran the value-based scorecard. The draft's opening paragraph carried one useful claim per forty words. The published version hit one per twelve. We didn't just trim fat—we reordered the entire argument so the cheapest backup tier appeared before the enterprise solution. That single shift changed which reader segment stayed engaged. The draft was written for IT managers. The edit aimed it at the freelancer who actually clicks publish.

The catch is that most editors stop at the surface. They count changes, not consequences.

Using a Simple Scorecard to Benchmark an Edit

Here's the three-column scorecard we used. Column one lists the reader's likely questions in order of urgency. Column two notes how quickly the draft answered each one—measured in words from the top. Column three does the same for the edited version. No fancy software. Just a spreadsheet and a willingness to be brutal about the draft.

Question one: "What's the cheapest way to back up 2TB?" The draft answered at word 410, after a history of tape drives. The edit answered at word 87. Question two: "How fast can I restore?" Draft: never explicitly, buried in a footnote. Edit: word 340, with a concrete example. Question three: "What's the catch with cloud storage?" Draft: word 720, vague. Edit: word 515, specific—egress fees, vendor lock-in, the math on a 200GB restore.

The scorecard exposes something ugly: a passage can be brilliant in isolation and still waste the reader's morning. We cut one genuinely clever analogy because it answered a question nobody had asked yet. That hurt—I liked the metaphor. But the benchmark didn't care about my ego.

'An edit is good when it shortens the distance from question to answer, not when it polishes every sentence into marble.'

— working note from the backup item's revision log, after we removed the second analogy

Real-World Example with Numbers

Let me give you the actual arithmetic. The draft scored 1.2 questions answered per 100 words. The published version scored 2.8. That's the headline metric—reader value density. But watch what happens when you slice it differently. For the freelancer segment, the edit reduced window-to-first-answer from 4.2 minutes to 1.1 minutes. For the IT manager segment? Barely moved—6.8 to 5.9 minutes. The edit was a targeted weapon, not a universal improvement.

What usually breaks first is the restore-window question. Few editors think to measure it because it lives outside the text. We had to instrument the page—tracking which anchor links people clicked, how far they scrolled before jumping to a new tab. That's the trade-off: value-based benchmarks demand you watch behavior, not just prose. Most CMS dashboards don't help. We built a crude tagger that flags any paragraph containing a question mark or a 'how'/'why' construction, then measures dwell window on those blocks.

The numbers surprised us. One section with zero questions had the highest engagement in the whole unit—a table comparing three backup providers side by side. So we adjusted the benchmark: tables and charts count as compressed answers. Pure prose isn't the only currency.

The honest limit here is sample size. One article doesn't prove anything. You need a dozen edits, each scored the same way, before patterns emerge. But even a single walkthrough like this—draft scored, edit scored, behavior tracked—reveals where your editing instincts are misaligned with reader needs. The first phase we ran it, we discovered our editor was polishing sentences that nobody read. Wrong order. The fix was not more editing. It was restructuring the component so the valuable parts appeared before the reader's patience ran dry.

Edge Cases: When the Metrics Get Tricky

Evergreen content vs. news stories

An evergreen component on setting up a home server rack can be edited today, revisited in six months, and still earn its keep. A breaking news item about a data-center outage? Its value decays by the hour. Value-based metrics that reward "window well spent" will naturally favor the evergreen component — and that's a problem if you're measuring an editor who primarily handles breaking news. You're not comparing like to like.

The fix isn't to abandon the metric. It's to calibrate the expected decay curve before you judge the edit. For news, a useful edit might mean cutting 200 words so readers get the core facts in under 90 seconds. For evergreen content, the same edit might mean adding 200 words of context that will still be accurate next year. Different goals, different denominators.

Most teams skip this. They track "reader retention" with one magic number and then wonder why their news editor looks like a slacker next to the person polishing SEO guides. Wrong order.

Niche audiences vs. broad reach

Here's where the metrics get genuinely sticky. A 3,000-word close look on Kubernetes security policies for a niche DevOps audience will have shorter reading times, fewer social shares, and lower raw engagement than a 600-word explainer on "what is a VPN" aimed at general readers. But the niche component might be doing far more valuable work — saving a senior engineer from a costly misconfiguration.

The catch is that value-based metrics are usually built on observable behavior: phase spent, scroll depth, repeat visits. None of those perfectly capture "this saved you from a bad decision." You can approximate it — follow-up searches on the same topic drop, comments show practitioners citing your framing — but it's always a proxy. I have seen editors gamify this by inflating word count on niche topics to hit a "time on page" target, which rewards padding. That's the opposite of good editing.

You'll need separate baselines per audience segment. Not one threshold for all content, but a range that shifts with the reader's existing expertise.

Long-form isn't inherently better; it's just easier to make look substantial. Short-form demands the discipline to say one thing well.

— observed from a technical editor who cut 40% of a tutorial without losing its core steps

Long-form vs. short-form: does length matter?

Length itself is the wrong variable. What matters is density of useful information per minute of reading. A 1,200-word piece that delivers one sharp, actionable insight beats a 3,000-word piece that repeats the same point in three different styles — even if the longer one keeps people on the page longer.

But the reverse is also true. Some topics require ramp-up. A complex legal analysis with multiple interacting clauses just can't be compressed to bullet points without losing nuance. The metric should reward matching the form to the content, not hitting a word count band.

The honest answer: you need per-format benchmarks, and you need to accept that comparing a 500-word product announcement to a 4,000-word investigative piece is apples-to-oranges. Set separate thresholds for each. Then watch for outliers — a short piece that holds readers for minutes, or a long piece that loses half its audience in the first two paragraphs. Those are the signals worth investigating.

The Honest Limits of Any Editing Benchmark

The danger of Goodhart's Law: when the metric becomes the target

Every benchmark I have ever designed eventually gets gamed. Not maliciously, usually—just by smart editors who notice that a specific score goes up when they do one specific thing. Change one word to boost "clarity" and the reader value score climbs. Do it a hundred times and you've got prose that reads like a manual for a toaster. That's Goodhart's Law in the wild: when a measure becomes the target, it stops being a measure.

The proposed metrics here are not immune. If you optimize hard for "reader retention per paragraph," you'll get short punchy paragraphs that never breathe. If you chase "vocabulary accessibility," you'll strip every word above eighth-grade level—including the ones that carry nuance. The catch is that any system you can articulate, someone can reverse-engineer.

What usually breaks first is the signal itself. Once editors know the formula, they write *for* the formula, and the metric no longer tracks quality—it tracks compliance. That's not a flaw in the math. It's a flaw in how we use numbers to replace judgment.

Measuring quality is inherently subjective

Let's be honest about the foundation. "Reader value" sounds objective, but the moment you define it, you're making a philosophical choice. One reader wants density; another wants momentum. One editor prizes surprise; another prizes clarity. The benchmark I outlined earlier assumes a certain kind of reader—one who stays, one who thinks, one who returns. That's a bet, not a fact.

I have watched two senior editors score the same paragraph, one giving it a 9 and the other a 3. Both were right. The paragraph was brilliant for a niche audience and opaque for everyone else. The metric didn't resolve the disagreement; it just gave each editor a number to argue with. That friction is real, and no formula removes it.

So when you implement these benchmarks, remember what you're actually measuring: *your* definition of value, frozen in time. It's a useful fiction, not a universal truth.

We can't automate taste

The final limit is the one nobody likes to admit. You can track reader behavior, word choice, sentence length, even emotional valence—but you can't automate the gut feeling that tells you a sentence is dead on arrival. Taste is built from thousands of small failures, and it lives in your nervous system, not in a spreadsheet.

Metrics are the flashlight, not the map. They show you where to look, not where to go.

— rule I repeat to every editor I train

That doesn't mean the flashlights are useless. It means you hold them in the hand that also knows how to feel for the walls. Use the benchmarks to surface problems, then override them with your own judgment when the context demands it. If the metric says "cut this aside," but the aside is the one thing that makes the piece human—cut the metric instead.

The honest limit is this: a benchmark is a mirror, and mirrors only show you what's already in front of you. They don't tell you what to wear. So build your dashboard, run the numbers, and then close the laptop when the real decision has to be made. You'll know the right call because it won't fit the formula.

Questions Editors Keep Asking

How do I get buy-in from skeptical writers?

Show them one article, edited the old way and the new way, side by side. Don't talk about metrics in the abstract — pull up a piece that got heavy line edits, then run it through a reader-value lens. Watch what happens when they see the edit that removed a confusing transition actually lift time-on-page by forty seconds. That changes the conversation from "you're grading my prose" to "we're both serving the reader." The catch is that you need one honest example before you get anyone on board, so pick a piece you edited yourself. Admit where the benchmark fails, too. Writers smell defensiveness instantly.

Start with a pilot group of two or three editors who trust you. Measure everything, but share only the top three numbers — clarity score, reread rate, and reader action. More than that gets noise. The tricky bit is that writer buy-in breaks when people feel surveilled, not measured. So make the benchmarks visible only to the editor until the writer asks. Then invite them into the dashboard. That reverses the power dynamic.

What tools should I use to measure these metrics?

You don't need a new platform. Most teams already have analytics, a CMS, and a readability check. We fixed this by building a simple spreadsheet that pulls time-on-page from Google Analytics, readability from Hemingway or Grammarly, and scroll depth from a small script. It took an afternoon to wire together. The fancy all-in-one tools exist, but they often measure proxy signals that don't match what you actually care about.

What usually breaks first is the scroll-depth script — ad blockers kill it, and mobile browsers throttle it. Have a fallback. I have seen teams abandon good benchmarks purely because the tooling felt incomplete, when a manual spot-check once a week would have been fine. Start ugly. Paper, a shared doc, a whiteboard. Refine later.

How often should I review benchmarks?

Monthly, but don't re-baseline everything each time. Review the top five articles from the previous month and compare their metric trajectories. Quarterly, look at the whole pool and adjust thresholds. Anything more frequent and you're chasing noise — a Tuesday holiday or a bot wave will wreck your confidence. That said, if a metric spikes or tanks by more than thirty percent, check it that week. Not because the metric is suddenly trustworthy, but because something in your editing process changed.

Don't these metrics just favor clickbait?

That's the question I hoped someone would ask — because a naive reader-value score absolutely rewards lurid headlines and listicle fluff. Reread rate filters some of that out: clickbait gets the click, then people leave fast. But not all of it. The honest limit is that any single metric can be gamed, which is why the benchmark includes a negative signal for discrepancy between headline promise and article content. Call it the "betrayal score." If the headline overpromises and the body underdelivers, the edit fails even when engagement looks great. It's not perfect — no benchmark is — but it catches the most common form of rot.

We don't measure whether an edit was good. We measure whether a reader's attention became something useful.

— private note I keep pinned above my desk, written after a week of chasing vanity metrics that felt great and taught us nothing.

Share this article:

Comments (0)

No comments yet. Be the first to comment!