Repository navigation
Usability Testing_2
Automated testing validates that Sentinel functions correctly from a technical perspective, but it does not validate whether an engineer under pressure can actually understand and operate the platform during a real incident.
The purpose of these usability tests was to evaluate whether a DevOps/SRE-oriented user can identify, understand, and resolve a production incident using Sentinel’s AI-assisted workflow without prior training.
Validate that a user can complete the core incident-response flow:
Detection → Analysis → Approval → Resolution
in less than 5 minutes and with enough confidence to make operational decisions safely.
If users cannot do this intuitively, the product fails to solve its primary operational problem.
Participant limitation: Due to academic timeline constraints, it was not possible to recruit only active on-call SRE/DevOps engineers. The tests were conducted with participants who partially represent the target audience through software engineering and infrastructure experience.
A total of 5 participants were recruited.
- Age: 21–24
- Familiar with:
- Docker
- Git
- basic cloud concepts
- No prior experience with:
- PagerDuty
- Grafana
- Datadog
- incident-management tools
Represents:
- junior DevOps onboarding experience
- Age: 25–32
- Experience with:
- deployments
- production troubleshooting
- container logs
- monitoring systems
Represents:
- partially experienced operational users
- In-person moderated sessions
- Live deployed Sentinel instance:
https://sentinel-softserve-1.onrender.com - Think-aloud protocol
Participants were instructed to verbalize their thoughts while interacting with the platform.
Moderators only intervened if participants became blocked for more than 2 minutes.
| Phase | Time |
|---|---|
| Introduction | 5 min |
| Task execution | 25–30 min |
| SUS + NPS questionnaire | 10 min |
| Open feedback | 5 min |
Average duration: 40–50 minutes
- Sentinel staging deployment
- Screen recording
- Observation notes
- SUS questionnaire
- NPS question
- Post-session interview
| Hypothesis | Result |
|---|---|
| H1 — Critical incidents can be identified quickly | Confirmed |
| H2 — Users understand the root cause without documentation | Confirmed |
| H3 — Approval flow generates confidence | Partially confirmed |
| H4 — Users can create incidents without training | Confirmed |
| H5 — Secondary flows are discoverable | Not confirmed |
The usability sessions aimed to answer the following questions:
Does the interface clearly communicate:
- where the user is?
- which incident is selected?
- current severity?
- current agent phase?
Is the AI reasoning panel:
- understandable?
- readable?
- excessively technical?
Can users explain the root cause without external help?
Before approving an action:
- do users understand the command?
- do they understand the risk?
- do they hesitate or approve automatically?
Does the incident timeline:
- tell a coherent operational story?
- or feel like disconnected logs?
Do sections such as:
- Labs
- post-mortem
- similar incidents
provide operational value or create unnecessary noise?
Tasks were goal-oriented instead of step-by-step instructions.
“You just started your on-call shift. Which incident would you handle first and why?”
Participant:
- identifies the most critical incident
- uses severity or operational context correctly
< 1 minute
- 5/5 participants completed successfully
- Average completion time: 18 seconds
Participants relied mainly on:
- severity colors
- the ⚡ badge
- incident status visibility
The incident list was consistently described as easy to scan.
“Explain what happened in app-demo and what the system proposes.”
Participant correctly identifies:
- incident type
- root cause
- at least one proposed action
< 2 minutes
- 5/5 participants completed successfully
- Average completion time: 1 min 50 sec
The AI reasoning panel was generally understandable.
However, junior participants struggled slightly with:
- operational terminology
- Labs terminology
“The system proposes restarting the container. Approve or reject the action.”
Participant:
- reads the command
- understands the risk
- makes a deliberate decision
< 1 minute
- 4/5 participants completed without assistance
- 1 participant required clarification
Average completion time: 55 seconds
The Approval Banner was highly visible and easy to notice.
However:
- 2 participants approved the action without fully reading the command
- both stated they trusted the AI automatically
This created an operational safety concern.
“Create an incident for a failure in the customers database.”
Participant successfully creates:
- title
- resource target
- severity
- saved incident
< 2 minutes
- 4/5 participants completed without assistance
- Average completion time: 2 min 30 sec
Participants described the creation flow as:
- straightforward
- clean
- predictable
The resource selector was consistently understood.
“Check whether this incident happened before and export its post-mortem.”
Participant:
- finds historical incidents
- accesses post-mortem
- exports report successfully
< 2 minutes
- 2/5 participants completed without help
- 2/5 required hints
- 1/5 failed
Average completion time: 3 min 10 sec
The Post-Mortem tab showed poor discoverability.
Most participants searched:
- in the incident header
- near action buttons
instead of using the tab navigation.
| Task | Without Help | With Hint | Failed |
|---|---|---|---|
| Quick triage | 5/5 (100%) | 0/5 | 0/5 |
| Understand incident | 5/5 (100%) | 0/5 | 0/5 |
| Approve action | 4/5 (80%) | 1/5 | 0/5 |
| Create incident | 4/5 (80%) | 1/5 | 0/5 |
| Export post-mortem | 2/5 (40%) | 2/5 | 1/5 |
77% without hints
Target defined in protocol: ≥ 80%
Result:
| Task | Target | Average |
|---|---|---|
| Quick triage | < 1 min | 18 sec |
| Understand incident | < 2 min | 1 min 50 sec |
| Approve action | < 1 min | 55 sec |
| Create incident | < 2 min | 2 min 30 sec |
| Export post-mortem | < 2 min | 3 min 10 sec |
| Participant | Profile | SUS Score |
|---|---|---|
| P1 | Student | 72.5 |
| P2 | Student | 67.5 |
| P3 | Student | 75.0 |
| P4 | Developer | 80.0 |
| P5 | Developer | 77.5 |
| Average | 74.5 / 100 |
A SUS score above 68 is considered above average.
Sentinel achieved:
74.5 / 100 → Good usability range
Question asked:
“Would you recommend this tool to a colleague?”
| Participant | Score |
|---|---|
| P1 | 7 |
| P2 | 6 |
| P3 | 8 |
| P4 | 9 |
| P5 | 8 |
7.6 / 10
Target defined in protocol: ≥ 7
Result: Target achieved
| Aspect | Average |
|---|---|
| Dashboard clarity | 4.4 / 5 |
| Understanding AI reasoning | 3.8 / 5 |
| Approval confidence | 4.0 / 5 |
| Overall ease of use | 4.1 / 5 |
All participants rapidly identified critical incidents using:
- severity colors
- badges
- status indicators
The dashboard scanning experience performed well under pressure.
The approval component was highly noticeable.
One participant stated:
“You can’t ignore it, which is exactly what you want at 2 AM.”
The creation flow was consistently completed with low friction.
Participants described it as:
- simple
- direct
- operationally focused
Tabbed navigation reduced cognitive overload.
Participants appreciated not having to scroll through long configuration sections.
High
3/5 participants could not naturally find the Post-Mortem section.
Add:
- visible “Post-Mortem” shortcut button
- contextual CTA when incident status is
resolved
High — Sprint 4
Medium
Junior participants struggled to understand the purpose of each Agent Lab.
Add:
- “View in action” links
- examples connected to real incidents
Medium — Sprint 4
Medium
2 participants approved commands without fully reading them.
Add:
- 2-second delay before enabling approval
- visible risk labels beside commands
Example:
- Low risk — restart container
- Medium risk — clear cache
- High risk — rollback database
Medium — Sprint 4
Low
The “Investigando” status was not visually distinct enough.
Add:
- subtle animation
- pulsing indicator
- processing spinner
Low — Sprint 4
Sentinel achieved:
- SUS: 74.5 / 100
- NPS: 7.6 / 10
- Task completion rate: 77% without assistance
The core operational workflow performed successfully:
Detection → Analysis → Approval
Participants were generally able to:
- identify incidents quickly
- understand the AI reasoning
- approve corrective actions confidently
This validates the main Sentinel design principle:
“Everything required to make a decision should be visible without excessive navigation.”
The primary usability problems were related to discoverability of secondary features rather than the critical operational workflow itself.
The identified improvements were prioritized for Sprint 4.
| Improvement | Priority | Sprint |
|---|---|---|
| Post-Mortem shortcut button | High | Sprint 4 |
| Approval safety delay + risk labels | Medium | Sprint 4 |
| Processing animation during investigation | Low | Sprint 4 |
| “View in action” Labs links | Low | Sprint 4 |
| Item | Value |
|---|---|
| Platform | In-person + screen recording |
| Environment | Sentinel staging deployment |
| Dataset | 10 seeded incidents |
| Session duration | 40–50 min |
| Execution period | Sprint 3 |
| Moderator | Nicol Garcia Tabares |
| Observer | Jacobo Montes |
- Seeded incidents prepared
- Test accounts created
- SUS questionnaire prepared
- Observation template prepared
- Sessions scheduled
- Consent obtained
- Same dataset used for all participants
Original protocol definition:
docs/sprint2-pruebas/03-protocolo-pruebas-usabilidad.md