Skip to content

Usability Testing_2

Nicolas Rico edited this page May 20, 2026 · 1 revision

Usability Testing — Sentinel

1. Purpose of the Usability Testing

Automated testing validates that Sentinel functions correctly from a technical perspective, but it does not validate whether an engineer under pressure can actually understand and operate the platform during a real incident.

The purpose of these usability tests was to evaluate whether a DevOps/SRE-oriented user can identify, understand, and resolve a production incident using Sentinel’s AI-assisted workflow without prior training.

Main Objective

Validate that a user can complete the core incident-response flow:

Detection → Analysis → Approval → Resolution

in less than 5 minutes and with enough confidence to make operational decisions safely.

If users cannot do this intuitively, the product fails to solve its primary operational problem.

Participant limitation: Due to academic timeline constraints, it was not possible to recruit only active on-call SRE/DevOps engineers. The tests were conducted with participants who partially represent the target audience through software engineering and infrastructure experience.


2. Participants

A total of 5 participants were recruited.

Profile A — Software Engineering Students (3 participants)

  • Age: 21–24
  • Familiar with:
    • Docker
    • Git
    • basic cloud concepts
  • No prior experience with:
    • PagerDuty
    • Grafana
    • Datadog
    • incident-management tools

Represents:

  • junior DevOps onboarding experience

Profile B — Developers with Production Experience (2 participants)

  • Age: 25–32
  • Experience with:
    • deployments
    • production troubleshooting
    • container logs
    • monitoring systems

Represents:

  • partially experienced operational users

3. Methodology

Session Format

  • In-person moderated sessions
  • Live deployed Sentinel instance: https://sentinel-softserve-1.onrender.com
  • Think-aloud protocol

Participants were instructed to verbalize their thoughts while interacting with the platform.

Moderators only intervened if participants became blocked for more than 2 minutes.


Session Duration

Phase Time
Introduction 5 min
Task execution 25–30 min
SUS + NPS questionnaire 10 min
Open feedback 5 min

Average duration: 40–50 minutes


Tools Used

  • Sentinel staging deployment
  • Screen recording
  • Observation notes
  • SUS questionnaire
  • NPS question
  • Post-session interview

4. Research Hypotheses

Hypothesis Result
H1 — Critical incidents can be identified quickly Confirmed
H2 — Users understand the root cause without documentation Confirmed
H3 — Approval flow generates confidence Partially confirmed
H4 — Users can create incidents without training Confirmed
H5 — Secondary flows are discoverable Not confirmed

5. Research Questions

The usability sessions aimed to answer the following questions:

Dashboard Understanding

Does the interface clearly communicate:

  • where the user is?
  • which incident is selected?
  • current severity?
  • current agent phase?

Agent Reasoning

Is the AI reasoning panel:

  • understandable?
  • readable?
  • excessively technical?

Can users explain the root cause without external help?


Approval Safety

Before approving an action:

  • do users understand the command?
  • do they understand the risk?
  • do they hesitate or approve automatically?

Timeline & Context

Does the incident timeline:

  • tell a coherent operational story?
  • or feel like disconnected logs?

Additional Features

Do sections such as:

  • Labs
  • post-mortem
  • similar incidents

provide operational value or create unnecessary noise?


6. Tasks and Results

Tasks were goal-oriented instead of step-by-step instructions.


Task 1 — Quick Triage

Scenario

“You just started your on-call shift. Which incident would you handle first and why?”

Success Criteria

Participant:

  • identifies the most critical incident
  • uses severity or operational context correctly

Target Time

< 1 minute

Results

  • 5/5 participants completed successfully
  • Average completion time: 18 seconds

Observations

Participants relied mainly on:

  • severity colors
  • the ⚡ badge
  • incident status visibility

The incident list was consistently described as easy to scan.


Task 2 — Understand the Incident

Scenario

“Explain what happened in app-demo and what the system proposes.”

Success Criteria

Participant correctly identifies:

  • incident type
  • root cause
  • at least one proposed action

Target Time

< 2 minutes

Results

  • 5/5 participants completed successfully
  • Average completion time: 1 min 50 sec

Observations

The AI reasoning panel was generally understandable.

However, junior participants struggled slightly with:

  • operational terminology
  • Labs terminology

Task 3 — Approve an Action

Scenario

“The system proposes restarting the container. Approve or reject the action.”

Success Criteria

Participant:

  • reads the command
  • understands the risk
  • makes a deliberate decision

Target Time

< 1 minute

Results

  • 4/5 participants completed without assistance
  • 1 participant required clarification

Average completion time: 55 seconds

Observations

The Approval Banner was highly visible and easy to notice.

However:

  • 2 participants approved the action without fully reading the command
  • both stated they trusted the AI automatically

This created an operational safety concern.


Task 4 — Create an Incident

Scenario

“Create an incident for a failure in the customers database.”

Success Criteria

Participant successfully creates:

  • title
  • resource target
  • severity
  • saved incident

Target Time

< 2 minutes

Results

  • 4/5 participants completed without assistance
  • Average completion time: 2 min 30 sec

Observations

Participants described the creation flow as:

  • straightforward
  • clean
  • predictable

The resource selector was consistently understood.


Task 5 — Export Post-Mortem

Scenario

“Check whether this incident happened before and export its post-mortem.”

Success Criteria

Participant:

  • finds historical incidents
  • accesses post-mortem
  • exports report successfully

Target Time

< 2 minutes

Results

  • 2/5 participants completed without help
  • 2/5 required hints
  • 1/5 failed

Average completion time: 3 min 10 sec

Observations

The Post-Mortem tab showed poor discoverability.

Most participants searched:

  • in the incident header
  • near action buttons

instead of using the tab navigation.


7. Evaluation Metrics

7.1 Task Success Rate

Task Without Help With Hint Failed
Quick triage 5/5 (100%) 0/5 0/5
Understand incident 5/5 (100%) 0/5 0/5
Approve action 4/5 (80%) 1/5 0/5
Create incident 4/5 (80%) 1/5 0/5
Export post-mortem 2/5 (40%) 2/5 1/5

Overall Completion Rate

77% without hints

Target defined in protocol: ≥ 80%

Result: ⚠️ Slightly below target


7.2 Average Task Time

Task Target Average
Quick triage < 1 min 18 sec
Understand incident < 2 min 1 min 50 sec
Approve action < 1 min 55 sec
Create incident < 2 min 2 min 30 sec
Export post-mortem < 2 min 3 min 10 sec

7.3 SUS — System Usability Scale

Participant Profile SUS Score
P1 Student 72.5
P2 Student 67.5
P3 Student 75.0
P4 Developer 80.0
P5 Developer 77.5
Average 74.5 / 100

Interpretation

A SUS score above 68 is considered above average.

Sentinel achieved:

74.5 / 100 → Good usability range


7.4 NPS — Net Promoter Score

Question asked:

“Would you recommend this tool to a colleague?”

Participant Score
P1 7
P2 6
P3 8
P4 9
P5 8

Average NPS

7.6 / 10

Target defined in protocol: ≥ 7

Result: Target achieved


7.5 Participant Satisfaction

Aspect Average
Dashboard clarity 4.4 / 5
Understanding AI reasoning 3.8 / 5
Approval confidence 4.0 / 5
Overall ease of use 4.1 / 5

8. Key Findings

Strengths Identified

Fast Incident Recognition

All participants rapidly identified critical incidents using:

  • severity colors
  • badges
  • status indicators

The dashboard scanning experience performed well under pressure.


Approval Banner Visibility

The approval component was highly noticeable.

One participant stated:

“You can’t ignore it, which is exactly what you want at 2 AM.”


Smooth Incident Creation

The creation flow was consistently completed with low friction.

Participants described it as:

  • simple
  • direct
  • operationally focused

Setup Navigation

Tabbed navigation reduced cognitive overload.

Participants appreciated not having to scroll through long configuration sections.


9. Problems Identified

P1 — Post-Mortem Discoverability

Severity

High

Problem

3/5 participants could not naturally find the Post-Mortem section.

Recommendation

Add:

  • visible “Post-Mortem” shortcut button
  • contextual CTA when incident status is resolved

Priority

High — Sprint 4


P2 — Labs Require More Context

Severity

Medium

Problem

Junior participants struggled to understand the purpose of each Agent Lab.

Recommendation

Add:

  • “View in action” links
  • examples connected to real incidents

Priority

Medium — Sprint 4


P3 — Blind Approval Risk

Severity

Medium

Problem

2 participants approved commands without fully reading them.

Recommendation

Add:

  • 2-second delay before enabling approval
  • visible risk labels beside commands

Example:

  • Low risk — restart container
  • Medium risk — clear cache
  • High risk — rollback database

Priority

Medium — Sprint 4


P4 — Investigating Status Visibility

Severity

Low

Problem

The “Investigando” status was not visually distinct enough.

Recommendation

Add:

  • subtle animation
  • pulsing indicator
  • processing spinner

Priority

Low — Sprint 4


10. Conclusions

Sentinel achieved:

  • SUS: 74.5 / 100
  • NPS: 7.6 / 10
  • Task completion rate: 77% without assistance

The core operational workflow performed successfully:

Detection → Analysis → Approval

Participants were generally able to:

  • identify incidents quickly
  • understand the AI reasoning
  • approve corrective actions confidently

This validates the main Sentinel design principle:

“Everything required to make a decision should be visible without excessive navigation.”

The primary usability problems were related to discoverability of secondary features rather than the critical operational workflow itself.

The identified improvements were prioritized for Sprint 4.


11. Planned Improvements

Improvement Priority Sprint
Post-Mortem shortcut button High Sprint 4
Approval safety delay + risk labels Medium Sprint 4
Processing animation during investigation Low Sprint 4
“View in action” Labs links Low Sprint 4

12. Session Logistics

Item Value
Platform In-person + screen recording
Environment Sentinel staging deployment
Dataset 10 seeded incidents
Session duration 40–50 min
Execution period Sprint 3
Moderator Nicol Garcia Tabares
Observer Jacobo Montes

13. Pre-Execution Checklist

  • Seeded incidents prepared
  • Test accounts created
  • SUS questionnaire prepared
  • Observation template prepared
  • Sessions scheduled
  • Consent obtained
  • Same dataset used for all participants

14. Reference

Original protocol definition:

docs/sprint2-pruebas/03-protocolo-pruebas-usabilidad.md

Clone this wiki locally