Keysight makes the instruments engineers use to test physical hardware: radios, chips, and the products built from them. Its KS8500B PathWave Test Automation Cloud runs those tests remotely across stationsA rack of test instruments wired together to run one set of tests. A lab holds many stations; a company may run labs on several continents. in labs worldwide, on a platform serving 300k+ engineers. Its dashboards are where teams watch system health, debug failures, and make time-sensitive calls.
I was the AI UX Researcher on the project, working with Keysight’s test-automation group in Silicon Valley from April to December 2025 alongside two other researchers. I moderated half the interviews, co-wrote the protocols, ran the relationship with Keysight, and led the design phase.
◇ marks a waypoint. Click one and the stage beside you changes.
Which stations are free, and why did last night’s run fail? Questions that simple moved across Excel, Slack, email, dashboards, and tools engineers had built themselves. Teams work in different time zones, so a missed signal stalls a test cycle until the next handoff.
The measurements were trustworthy; the path from them to a decision was not. We read six papers, ran a , and turned the brief into that every method after them was chosen to answer.
Seven hour-long interviews with lab managers, test engineers, and design engineers mapped how work actually gets done. Four think-aloudThe participant works through real tasks while narrating their thinking out loud, so you hear where expectations break rather than guessing from the result. usability tests, ten tasks each, ran on and showed where those workflows broke.
A System Usability ScaleSystem Usability Scale: ten standard agree-or-disagree statements that produce a single comparable usability score. survey set a baseline, and a skip-logicA survey that branches: the next question depends on your last answer, so people only see what applies to them. survey at a Keysight event in Denmark tested those patterns with fifty more people. the answer was the same: the data existed, but nothing said what needed attention or what to do next.
Several participants stood in for lab managers rather than being one, and the findings say so. It changed who we recruited: we collected consented contacts at the event, and the final round of testing ran with real test engineers and a bench lead.
Every transcript was coded twice: once for what was saidA short label for what a line of transcript is about, written close to the participant's own words., then again for what those labels had in commonAxial coding: grouping those first labels into larger categories, so scattered remarks become a pattern.. Those codes were clustered on an against the research questions, and each usability session was drawn as a .
From 900+ rows came four themes: Visibility for decision-making, Understandability, Workflow efficiency, and Customizability. Two artifacts carried them into design: a and .
Mid-project Keysight moved the dashboards onto Apache Superset, so before drawing we ran a of the new foundation, rewrote the against what it found, under the four themes, and mapped the .
The first buildable concept was one dashboard assembled from role templates: pick the template for your job, adjust it, see everything in a single place. It answered the research directly, and it was cheap to test.
We ran on the wireframes. Participants could work the template flow, but they already lived in Grafana and TableauThe two dashboard tools most of these engineers already used daily: Grafana for live monitoring, Tableau for analysis.. A Keysight dashboard that looked like either gave them no reason to switch, and one that worked differently would fight habits they had spent years building. Jakob’s law names the trap.
How do we make this meaningfully better than the dashboards users already trust?
AI was the obvious answer, and Keysight pushed back on it: were we adding AI for its own sake? So we made that the test. We built both answers, and , put both buttons on one screen, and asked every participant why they reached for the one they did.
, out of a general distrust of AI, then preferred the AI path once they had tried it. Their worry was never the output: it was what the AI would read, and where that data would go. The shipped direction keeps both, and its three features each shorten one stretch of the path from data to decision.
Creation starts from a dataset and a sentence. You pick the data, describe what you need in plain language, the AI proposes charts, and you drag the useful ones onto the canvas. carry the detail work.
: suggestions beside the canvas, never instead of it. They make starting fast while leaving engineers in control of what their team will rely on. The data-access step and the terms of use came straight from the trust finding.
The main view combines filters, , and an AI summary of what the dashboard currently shows. The charts stay plain and legible. At scale, .
The first version hid summaries inside each chart, visible only when expanded. : anything unusual detected automatically and shown in one dedicated place. The summary moved to the left panel, always visible, with its top three insights each linking to the chart behind them. That link also answered the second round’s request to check the AI’s claims against the data.
Why put AI in the summary and not in the charts?
The gap was interpretation, not visualization. Engineers trusted their charts and knew how to read them. What they lacked was a fast answer to which signal needs attention, and the summary answers exactly that.
Drill-downsClicking into a chart to see the finer-grained data behind a number, then the records behind that. trace a signal to its cause without leaving the tool, and from any analysis the AI drafts a report. Writing findings up elsewhere was the last forced exit from the tool, and the drafts remove it.
: brief on screen, with the detail arriving in the downloaded file. The final design added an assistant for revising the draft, insights pulled from the AI summary, and share, export, and save-as-template. makes experimenting safe on a dashboard a whole team depends on.
Nine think-aloud sessions shaped the designs: five on wireframes, four on the high-fidelity prototype with test engineers and a bench lead. A separate timed study measured them. Every participant ran both dashboardsWithin-subject: every participant runs both versions, so each person is their own comparison and individual differences cancel out., the old one first and the AI-assisted one second, on the same analysis and drill-down tasks.
The number is conservative, and it is narrow. Creating dashboards and writing reports were never timed, and those are the stretches the AI shortens most. The prototype ran on simulated data with the AI loosely connected, so people still checked its output by hand. The next study is the same tasks, in production, on real data.
What the evidence cannot say yet