Published August 2026 | Version v1
Dissertation Embargoed

Robust Multi-Pattern Detection of Disk Fail-Slow Behaviors at Scale

Creators

  • 1. ROR icon University of Chicago

Contributors

Committee members:

  • 1. ROR icon University of Chicago
  • 2. ROR icon Shanghai Jiao Tong University

Description

Large-scale cloud systems depend on predictable performance from large fleets of servers and storage devices. Yet many harmful failures do not appear as clean crashes. Disks may continue to serve requests while becoming abnormally slow, and servers may remain available while delivering degraded benchmark performance. These fail-slow and underperformance cases are difficult to detect because production telemetry is noisy, heterogeneous, and often sparse. Removing healthy components wastes fleet capacity, while missing degraded components can increase tail latency, reduce efficiency, and create costly operational work.

This dissertation argues that reliable performance-degradation detection must be both pattern- aware and data-aware. First, it presents a large-scale study of disk fail-slow behavior using production data covering over 55 million drive days. The study shows that fail-slow disks do not follow a single anomaly shape: they appear in five major spatial and temporal patterns, ranging from obvious outliers to subtle cases that overlap with normal behavior. It also identifies noisy patterns that explain why simple thresholds and prior detectors miss important cases or generate excessive false positives.

Second, the dissertation presents KRATOS, a robust multi-pattern disk fail-slow detection method. KRATOS combines a lightweight temporal feature, divergence-based risk scoring, and denoising mechanisms for difficult in-between cases. Evaluated against nine classic and state-of-the-art methods across production, public, independent deployment, and fault-injection datasets, KRATOS is the only approach that detects all observed fail-slow patterns while maintaining stable performance across environments.

Finally, the dissertation extends this philosophy from disk telemetry to server-level benchmark data. It presents FREP, a fleet performance evaluation pipeline for identifying underperforming servers from more than 3 million memory benchmark observations across 1 million production servers. FREP uses feature analysis to form meaningful configuration-matched cohorts, derives conservative performance floors from stable peers, and applies statistical testing with false-

discovery control. Compared with six standard unsupervised anomaly detection methods, FREP achieves near-perfect precision while reducing operator burden by orders of magnitude.

Together, these contributions show that effective fleet diagnosis cannot treat performance data as a generic anomaly-detection problem. Robust detection requires first understanding how degradation manifests, then building baselines and alerting mechanisms that respect the structure, noise, and sampling constraints of production data.

Files

Embargoed

The files will be made publicly available on August 22, 2028.

Additional details

UChicago Information

Division(s)
Physical Sciences Division
Department(s)
Computer Science