Short Scales, Wide Perspectives: Bridging the Accessibility Gap with SUS and UMUX

Short Scales of Satisfaction Assessment: A Proxy to Involve Disabled Users in the Usability Testing of Websites

2015-01-01
Simone Borsci, Stefano Federici, Maria Laura Mele, Matilde Conti
Summary
Problem
Method
Results
Takeaways
Abstract

This study evaluates the psychometric reliability of three short satisfaction scales—SUS (System Usability Scale), UMUX (Usability Metric for User Experience), and UMUX-LITE—when applied to blind users. By comparing blind and sighted participants navigating a public transportation website, the research demonstrates that these compact tools effectively capture the divergent usability experiences of disabled users.

TL;DR

In the world of UX, "fast and cheap" usually means "less accurate." However, research by Borsci et al. suggests that ultra-short satisfaction scales like the System Usability Scale (SUS) and UMUX-LITE are highly reliable tools for capturing the unique frustrations of blind users. By including a small sample of disabled users and using these scales, practitioners can expose critical accessibility failures that sighted users never encounter—all without blowing the research budget.

Background: The High Cost of Exclusion

Excluding disabled users from usability testing is often justified by practitioners as a "cost-saving measure." Testing with blind users or those with motor impairments frequently involves assistive technologies, longer sessions, and specialized protocols.

The author's core insight is that satisfaction—one of the three pillars of usability (alongside effectiveness and efficiency)—can be measured rapidly using existing short-form questionnaires. If these scales are proven reliable for disabled cohorts, they serve as a "stress test" proxy, allowing teams to integrate diverse perspectives with minimal friction.

Problem & Motivation: Are the Scales Universal?

While the SUS (10 items) and the newer UMUX (4 items) and UMUX-LITE (2 items) are industry standards, their psychometric properties had not been rigorously tested with disabled populations. Does a blind user perceive a question like "I found the system unnecessarily complex" the same way a sighted user does? If the reliability (Cronbach’s ) holds up, then these scales can be used to compare "apples to apples" across different user groups.

Methodology: The Trenitalia Stress Test

The researchers recruited 10 blind and 10 sighted users to perform three common tasks on the Italian public train website (booking a ticket, finding an info-point, and filing a claim).

The Protocol:

  1. Interaction: Users navigated the site using "thinking aloud" protocols.
  2. Assessment: After the tasks, users filled out SUS, UMUX, and UMUX-LITE in random order.
  3. Analysis: The team compared the number of errors found and the statistical reliability of the scales.

UMUX-LITE Adjustment Formula The study utilized the standard regression formula to align UMUX-LITE scores with the SUS 0-100 range.

Key Results: Two Different Worlds

The experiment revealed a stark contrast between user groups:

  • Error Detection: 29 total problems were found. Blind users found 19; sighted users only 10. 11 problems were unique to blind users, meaning a sighted-only test would have missed over 50% of the accessibility blockers.
  • Satisfaction Scores: Sighted users gave the site a "C" grade. Blind users gave it an "F."

Scale Reliability

The statistics confirmed that these short questionnaires are robust:

  • SUS and UMUX-LITE showed high internal consistency (Cronbach’s ) for both groups.
  • UMUX (4-item version) struggled with the blind cohort (α = 0.568), suggesting that the negative-toned items in the 4-item version might be confusing or perceived differently by screen-reader users.

Summary of Comparison Results Note the drastic drop in SUS scores from sighted (67.75) to blind (15.25) users.

Critical Insight: The Power of Mixed Cohorts

The most striking takeaway is the "Aggregated Score." When you combine the feedback of both groups, a website that looks "average" (Grade C) to a sighted user is revealed to be "failing" (Grade F).

By using SUS or the hyper-efficient UMUX-LITE (just 2 questions!), UX practitioners can:

  1. Quantify the Accessibility Gap: Show stakeholders exactly how much harder the experience is for disabled users in a single number.
  2. Save Costs: Use these scales as a quick "pulse check" in remote or unmoderated tests with disabled users.
  3. Standardize Reporting: Use the same metrics for all users, making the inclusive design data easy to digest alongside general usability data.

Conclusion & Future Work

The study concludes that SUS and UMUX-LITE are reliable proxies for involving disabled users in usability testing. While the 4-item UMUX requires more investigation regarding its phrasing for blind users, the 2-item UMUX-LITE emerges as a surprisingly strong alternative. To build truly universal public services, we must stop viewing disabled users as "edge cases" and start using these efficient tools to bring their divergent perspectives into the heart of the design process.

Find Similar Papers

Try Our Examples

  • Search for recent studies comparing the reliability of UMUX-LITE versus SUS in accessibility-focused web audits.
  • Which paper first established the UMUX-LITE formula for score adjustment to match SUS, and has it been refined for specific disability types?
  • Explore how these short satisfaction scales have been applied to mobile application usability testing for users with motor impairments or cognitive disabilities.
Contents
Short Scales, Wide Perspectives: Bridging the Accessibility Gap with SUS and UMUX
1. TL;DR
2. Background: The High Cost of Exclusion
3. Problem & Motivation: Are the Scales Universal?
4. Methodology: The Trenitalia Stress Test
4.1. The Protocol:
5. Key Results: Two Different Worlds
5.1. Scale Reliability
6. Critical Insight: The Power of Mixed Cohorts
7. Conclusion & Future Work