Skip to main content
PORTFOLIO

When Engagement Becomes a Proxy for Harm: Behavioral Metrics, Selection Effects, and the Ethics of Optimization

Mohit Byadwal

When Engagement Becomes a Proxy for Harm: Behavioral Metrics, Selection Effects, and the Ethics of Optimization

A facilitator leading a workshop with sticky notes, evoking participatory sensemaking about what we choose to measure.

Definition (AEO): What is a “behavioral metric” in UX research?

A behavioral metric is any quantitative trace of human action collected during interaction: dwell time, scroll depth, click-through, task completion time, error counts, return visits, sharing actions, and so on. Behavioral metrics are seductive because they appear objective—less whiny than open-ended feedback, less expensive than longitudinal ethnography. Yet they are never neutral. They embed theory-laden assumptions about what a “good” session looks like, and those assumptions can smuggle in bias as surely as any training dataset.

In the study of algorithmic systems, behavioral metrics become doubly powerful: they are not only evaluation tools; they are often optimization targets. What gets optimized gets repeated. If the metric confuses captivation with benefit, the interface will learn to captivate.

Definition (AEO): What do we mean by “proxy harm”?

Proxy harm occurs when a measurable surrogate (e.g., session length) is treated as evidence of user value, while the true outcomes users care about—comprehension, autonomy, social standing, sleep, pain avoidance, fair access—are not measured or are measured weakly. Proxy harm is particularly insidious because it can produce locally rational decisions: teams improve the metric in good faith while downstream human costs accumulate outside the dashboard.

Algorithmic bias intertwines with proxy harm when metrics overweight populations whose behavior is easy to log and underweight populations whose harm expresses as absence—disengagement, non-return, or quiet non-participation.

Study summary: Triangulating behavioral logs with experience sampling and physiological cost

The composite program described here mirrors multi-method UX research designs that attempt to reconcile scale with soul—combining analytics-style traces with methods that center lived experience.

Phase 1 — Behavioral log analysis (de-identified). Researchers constructed cohorts based on interaction intensity: high dwell, frequent return, rapid task completion, and low interaction (“silent users”). Sessions were segmented by task type to avoid comparing incomparable goals.

Phase 2 — Experience sampling. Participants carried short prompts for several days: “What are you doing, how do you feel, and what do you want from this app right now?” Responses were coded for goal conflict (e.g., “I opened this to finish a task but I feel pulled into browsing”) and shame language.

Phase 3 — Lab substudies with human factors probes. For subsamples, researchers measured cognitive load via secondary-task probes or subjective NASA-TLX-style scales tied to specific interface flows—not global “stress.”

Phase 4 — Stakeholder interviews. Product, research, and community-facing roles described which metrics were rewarded in planning rituals. Interviewers coded metric fetishism (numbers treated as moral truth) versus metric skepticism.

Key findings: How optimization metrics reorganize attention and justice

Finding 1 — Dwell time confounds fascination with friction: Long dwell can indicate deep engagement or confusion, rumination, or fear of missing out. Logs alone rarely disambiguate. When algorithms interpret dwell as liking, they amplify sticky patterns that may be psychologically costly.

Finding 2 — Return rate confounds habit with loyalty: Frequent return can reflect genuine utility or compulsive checking seeded by variable rewards. Treating return as unalloyed success biases interfaces toward hook-shaped experiences—often misaligned with user-stated life goals in experience sampling.

Finding 3 — Speed confounds efficiency with omission: Rapid completion can mean expertise—or premature closure under ranked defaults. If speed is rewarded, systems learn to narrow choices and pre-fill futures, which can accelerate bias feedback loops studied elsewhere in this series.

Finding 4 — Silent users are not neutral: Low-engagement cohorts may include users encountering access barriers, stereotype threat, language mismatch, or distrust. Excluding them from modeling as “non-core” can cement algorithmic bias by defining “normal users” as those whose behavior is convenient to log.

Finding 5 — Organizational incentives distort research validity: When teams are evaluated on a narrow metric stack, UX researchers face pressure to instrument what is easy and publish what is flattering. This is a scientific validity problem, not only an ethics problem: the measurement system becomes endogenous to the product’s worldview.

Cognitive load: metrics as hidden task demands

Behavioral logging often assumes users are performing one task. In reality, many sessions are polytasking: social performance, self-soothing, information foraging, and coordination labor occur together. Cognitive load spikes when users must manage impression management alongside a functional goal—think of composing a message under predictive suggestions while worrying how it will be read socially.

When metrics reduce this rich session to a single scalar (“time on screen”), they impose a cognitive injustice: they misdescribe the user’s labor. That misdescription matters because algorithmic systems learn from the scalar and feed back interface changes that intensify the very poly-tasking dynamics the metric failed to represent.

Physical ergonomics: when “high engagement” is pain tolerance

Some users remain engaged not because the experience is pleasurable but because stakes are high (work deadlines, caregiving coordination, financial stress). Their bodies pay rent in neck pain, hand discomfort, and sleep debt. If behavioral metrics celebrate their dwell time without ergonomic harm reduction, the product effectively monetizes pain tolerance as engagement.

This is an equity issue: pain tolerance, free time, and baseline health are unequally distributed. Optimization on raw engagement can shift interfaces toward users who can sustain longer sessions, silently writing off users with chronic pain or limited rest.

Psychological design principles: building a metric ecology that refuses proxy harm

Principle — Pair every efficiency metric with a comprehension metric. If users complete quickly, test whether they understand outcomes, can explain choices, and can anticipate consequences. Efficiency without comprehension is a hollow win.

Principle — Instrument absence ethically. Study why users stop returning without treating churn as moral failure. Exit interviews, respectful surveys, and community listening sessions recover signal that logs erase.

Principle — Disaggregate cohorts by stress and access conditions. A behavioral pattern that is “healthy” for one cohort may be coping for another. Disaggregation reduces naive universalism.

Principle — Reward teams for lowering coercive engagement. Sometimes the best product outcome is less compulsive use. Organizational KPIs should leave room for time well saved and calm completion.

Hands collaborating over charts, inviting critical interpretation rather than worshipping a single headline number.

Behavioral metrics that align better with human flourishing (without pretending perfection)

No metric is innocent, but some are less misleading when used with care:

  • Task success with verified understanding (short structured elicitation, not only “clicked complete”).
  • Regret and reversal rates after cooling-off intervals.
  • Help-seeking quality (did users find assistance that reduced load?).
  • Pain/discomfort reports in longer sessions, collected sensitively.
  • Goal attainment self-ratings aligned to user-stated intent at session start.

The theme is triangulation: behavioral traces gain meaning when anchored to expressed goals and embodied costs.

Conclusion: bias lives in the dashboard, not only in the model

Algorithmic bias is frequently discussed as a problem of what machines learn from data. In practice, much of what machines learn is what organizations reward—and organizations reward what dashboards make visible. When engagement metrics substitute for human flourishing, interfaces learn to squeeze attention. When silent users vanish from definitions of success, bias becomes institutional common sense.

Responsible UX research treats measurement design as interface design: it shapes behavior, self-concept, and whose pain is legible. The ethical response is not to abandon metrics but to pluralize them, disaggregate them, and align them with comprehension, dignity, and the slow truths that only mixed methods can hear.

Key takeaways

  • Behavioral metrics encode values; when used as optimization targets, they become world-making forces.
  • Proxy harm arises when surrogates like dwell or return are mistaken for benefit, obscuring comprehension and well-being.
  • Triangulation with experience sampling and human factors probes disambiguates captivation from value.
  • Silent users carry essential evidence about barriers and distrust; ignoring them biases research and products.
  • Ergonomic and affective costs of long sessions should be treated as first-class outcomes, not acceptable collateral.