Robust and Rapid Clustering of KPIs for Large-Scale Anomaly Detection

Zhihan Li, Youjian Zhao, Rong Liu, Dan Pei

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

69 Scopus citations

Abstract

For large Internet companies, it is very important to monitor a large number of KPIs (Key Performance Indicators) and detect anomalies to ensure the service quality and reliability. However, large-scale anomaly detection on millions of KPIs is very challenging due to the large overhead of model selection, parameter tuning, model training, or labeling. In this paper we argue that KPI clustering can help: we can cluster millions of KPIs into a small number of clusters and then select and train model on a per-cluster basis. However, KPI clustering faces new challenges that are not present in classic time series clustering: KPIs are typically much longer than other time series, and noises, anomalies, phase shifts and amplitude differences often change the shape of KPIs and mislead the clustering algorithm. To tackle the above challenges, in this paper we propose a robust and rapid KPI clustering algorithm, ROCKA. It consists of four steps: preprocessing, baseline extraction, clustering and assignment. These techniques help group KPIs according to their underlying shapes with high accuracy and efficiency. Our evaluation using real-world KPIs shows that ROCKA gets F-score higher than 0.85, and reduces model training time of a state-of-the-art anomaly detection algorithm by 90%, with only 15% performance loss.

Original languageEnglish
Title of host publication2018 IEEE/ACM 26th International Symposium on Quality of Service, IWQoS 2018
ISBN (Electronic)9781538625422
DOIs
StatePublished - 22 Jan 2019
Event26th IEEE/ACM International Symposium on Quality of Service, IWQoS 2018 - Banff, Canada
Duration: 4 Jun 20186 Jun 2018

Publication series

Name2018 IEEE/ACM 26th International Symposium on Quality of Service, IWQoS 2018

Conference

Conference26th IEEE/ACM International Symposium on Quality of Service, IWQoS 2018
Country/TerritoryCanada
CityBanff
Period4/06/186/06/18

Fingerprint

Dive into the research topics of 'Robust and Rapid Clustering of KPIs for Large-Scale Anomaly Detection'. Together they form a unique fingerprint.

Cite this