This blog relates to experiences in the Systems Capacity and Availability areas, focusing on statistical filtering and pattern recognition and BI analysis and reporting techniques (SPC, APC, MASF, 6-SIGMA, SEDS/SETDS and other)
Popular Post
-
I have got the comment on my previous post “ BIRT based Control Chart “ with questions about how actually in BIRT the data are prepared for ...
-
Re-posting interesting article from R-vs.-Python-for-Data-Science R vs. Python for Data Science Norm Matloff, Prof. of Computer ...
_
Monday, February 27, 2017
CMG'12 Trubin's 2nd Presentation : SEDS-Lite
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Saturday, January 7, 2017
US Patent "SYSTEMS AND METHODS FOR MODELING COMPUTER RESOURCE METRICS", I Trubin et al (granted as #10,437,697)
UPDATE 10/15/2019 - Patent is granted. The Number is 10,437,697
SYSTEMS AND METHODS FOR MODELING COMPUTER RESOURCE METRICS
I Trubin, M Schutt, J Robinson - United States Patent Application 15/184501
This disclosure relates generally to system modeling, and more particularly to systems and methods for modeling computer resource metrics. In one embodiment, a processor-implemented computer resource metric modeling method is disclosed. The method may include detecting one or more statistical trends in aggregated interaction data for one or more interaction types, and mapping each interaction type to one or more devices facilitating the transactions. The method may further include generating one or more linear regression models of a relationship between device utilization and interaction volume, and calculating one or more diagnostic statistics for the one or more linear regression models. A subset of the linear regression models may be filtered out based on the one or more diagnostic statistics. One or more forecasts may be generated using the remaining linear regression models, using which a report may be generated and provided.
____
NB: This patent application uses statistical exception and trend detection (SETDS) data to do modeling.
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Tuesday, December 27, 2016
How Bayesian inference works. Can that improve the SPC?
What about using that approach for SPC? At least one article is HERE about it:
Bayesian Statistical Process Control
and some dispute is HERE:
Why isn't bayesian statistics more popular for statistical process control
Note there is a "Bayesian statistical process control chart"...
In my experience we often do not have too many data to be used for mean, UCL and LCL calculations, so based on "An Application to Bayesian Methods in SPC":
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Friday, December 23, 2016
CMG Amplify - new blog about Capacity and Performance
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Wednesday, December 7, 2016
The BANK should be a tech company to win the market
I work for Capital One bank as IT SME and naturally I support the direction it goes. Friends keep asking me how the bank could be a tech company like Netflix? The best who can answer that question is the CIO of Capital One. HERE IT IS:
Capital One rides the cloud to tech company transformation
I finally am getting use to a new style of workplace I have now:![]() |
| The picture from the article linked above |
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Friday, December 2, 2016
CMG'16 (#imPACt) aftershocks: Could we WAC the #Cloud? or how to build cube of cubes
One of the finding from my CMG'16 conference attending and speaking with the vendor's representatives is the 6fusion way to measure systems overall status by WAC - Workload Allocation Cube.
"The tool measures the resource consumption of AWS compute (EC2) instances, along with Elastic Block Storage (EBS) volumes. It does this via 6fusion's "Workload Allocation Cube" (WAC) technology, which works by measuring datapoints that include CPU utilization, disk utilization, storage capacity, and disk, WAN, and LAN IOPS. This information is aggregated through its WAC technology to output a single value that reflects the performance and resource use of an app." - http://www.theregister.co.uk/2013/07/29/6fusion_workload_allocation_cube/
My comment so far is following:
Long ago I tried to use "System Health Index" (from Concord eHeallth performance monitor) to estimate the system usage based on 5 main subsystems measures and published my thoughts about it in my CMG'03;06;07 papers; E.G. see some details in the following post.
"The tool measures the resource consumption of AWS compute (EC2) instances, along with Elastic Block Storage (EBS) volumes. It does this via 6fusion's "Workload Allocation Cube" (WAC) technology, which works by measuring datapoints that include CPU utilization, disk utilization, storage capacity, and disk, WAN, and LAN IOPS. This information is aggregated through its WAC technology to output a single value that reflects the performance and resource use of an app." - http://www.theregister.co.uk/2013/07/29/6fusion_workload_allocation_cube/
My comment so far is following:
Long ago I tried to use "System Health Index" (from Concord eHeallth performance monitor) to estimate the system usage based on 5 main subsystems measures and published my thoughts about it in my CMG'03;06;07 papers; E.G. see some details in the following post.
Disk Subsystem Capacity Management - my CMG'03 paper - "Health Index" metric and Dynamic Thresholds
Having that metric recorded I have suggested (in my CMG'07 paper) to use the Tree-map report against that to get one overall health check picture of numerous systems.
Also I have applied my anomaly detection technique (SEDS) against that Health Index metric to detect at once any abnormal cases across all main 5 subsystems:
I have published a few examples in my CMG'06 SYSTEM MANAGEMENT BY EXCEPTION, PART 6
Anyway I see some similarity between modern "WAC" and the olde good "Health Index". Do you see it as well?
P.S. One more thing.... WAC is a cube and the SEDS data (SEDS profile) in the picture above is also a data cube that represent a "signature" of the object (server in the case). Actually if the Health index a cube also, the SEDS Health index profile would be a cube of cubes....!? ,
Anyway here is the discussion how that SEDS profile data cube can be built using the open source tools (my most visited post in this blog, btw...):
P.S. One more thing.... WAC is a cube and the SEDS data (SEDS profile) in the picture above is also a data cube that represent a "signature" of the object (server in the case). Actually if the Health index a cube also, the SEDS Health index profile would be a cube of cubes....!? ,
Anyway here is the discussion how that SEDS profile data cube can be built using the open source tools (my most visited post in this blog, btw...):
One Example of BIRT Data Cubes Usage for Performance Data Analysis
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Tuesday, November 29, 2016
Invitation to the Advanced Software in #Robotics conference in LIEGE (Belgium-1983) to present my paper
I was a co-author of the paper "Mathematical simulation of tasks of robot operation accuracy and readability". See the session 4 in the agenda above. (Note my initial is misspelled as C.A. TRUBIN , should be I.A. TRUBIN).

That was USSR time and in spite we were invited and even sent our paper translated into English, instead of us some communist functioner went there and even did not appear at the conference... So it was not really published...
I have found the online documents so far related to this - http://ieeexplore.ieee.org/document/4336367/
- http://dl.acm.org/citation.cfm?id=577664
- https://www.amazon.com/Advanced-Software-Robotics-International-Proceedings/dp/0444868143 (where the preceedings could be bought!)
That was USSR time and in spite we were invited and even sent our paper translated into English, instead of us some communist functioner went there and even did not appear at the conference... So it was not really published...
I have found the online documents so far related to this - http://ieeexplore.ieee.org/document/4336367/
- http://dl.acm.org/citation.cfm?id=577664
- https://www.amazon.com/Advanced-Software-Robotics-International-Proceedings/dp/0444868143 (where the preceedings could be bought!)
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Saturday, November 12, 2016
Help us generate content for the CMG blog -LinkedIn CMG group discussion
https://www.linkedin.com/groups/125056/125056-6202640842222641155
Be the hero of CMG, write a blog post! - Renato Bonomini
At the #CMGimPACt conference, did you learn something that you'd like to share? - Todd Minnella
Share your learnings, they'll make a great blog post! Did you leave with more questions than answers? Share your questions, they will be our content ideas!
Do you have a crazy idea? Let us know! - Melanie Heimer
Share your ideas in the comments, or contact me or Igor Trubin directly!
Be the hero of CMG, write a blog post! - Renato Bonomini
At the #CMGimPACt conference, did you learn something that you'd like to share? - Todd Minnella
Share your learnings, they'll make a great blog post! Did you leave with more questions than answers? Share your questions, they will be our content ideas!
Do you have a crazy idea? Let us know! - Melanie Heimer
Share your ideas in the comments, or contact me or Igor Trubin directly!
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Monday, November 7, 2016
Me presenting at CMG.org conference "Is your Capacity Available?"
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Sunday, November 6, 2016
Sitting on the Board of Directors (www.CMG.org)
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
#CMGnews - "PERFORMANCE OR CAPACITY? DIFFERENT APPROACHES FOR DIFFERENT TASKS" by #Facebook SMEs
I am at CMG'2016 conference (imPACt) right now and going to attend the following session:
PERFORMANCE OR CAPACITY?
DIFFERENT APPROACHES FOR DIFFERENT TASKS
by Alexander Gilgur, Steve Politis (Facebook, Inc.)
ABSTRACT
We often talk about performance and capacity as one thing, and indeed they complement each other in a powerful balancing loop: higher capacity improves performance, decreased performance indicates insufficient capacity, which needs to be provisioned for. However, we often miss the fact that measurement and aggregation approaches that are used in performance monitoring are not always useful for capacity planning, while approaches that we use in capacity planning are often meaningless for performance analysis. This paper explores this gap and discusses ways to reconcile the two tasks.
I really appreciate they have mentioned my following work in their paper:
Trubin, I. (2006) System Management by Exception. Presented at the Annual International Conference of the Computer Measurement Group (CMG 2006). Reno, NV. December 2006.
But I feel like my later publication was more relevant to the subject:
PERFORMANCE OR CAPACITY?
DIFFERENT APPROACHES FOR DIFFERENT TASKS
by Alexander Gilgur, Steve Politis (Facebook, Inc.)
ABSTRACT
We often talk about performance and capacity as one thing, and indeed they complement each other in a powerful balancing loop: higher capacity improves performance, decreased performance indicates insufficient capacity, which needs to be provisioned for. However, we often miss the fact that measurement and aggregation approaches that are used in performance monitoring are not always useful for capacity planning, while approaches that we use in capacity planning are often meaningless for performance analysis. This paper explores this gap and discusses ways to reconcile the two tasks.
I really appreciate they have mentioned my following work in their paper:
Trubin, I. (2006) System Management by Exception. Presented at the Annual International Conference of the Computer Measurement Group (CMG 2006). Reno, NV. December 2006.
But I feel like my later publication was more relevant to the subject:
Exception Based Modeling and Forecasting - CMG'2008
Abstract
How often does the need arises for modeling and forecasting? Should it be done manually by ad-hoc, by project requests or automatically? What tools and techniques are best for that? When is trending forecast enough and when is a correlation with business drivers required? The answers to these questions are presented in this session. The capacity management system should automatically provide a small list of resources that needs to be modeled or forecasted; a simple spreadsheet tool can be used for that. This technique method is already implemented on the author’s environment with thousands of servers.
I am looking forward meeting the presenter (my long time CMG friend Alex) to discuss this....
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Tuesday, November 1, 2016
Southern CMG meeting presentations are published (#AppDynamics, #Cibra, #IBM, #DataKinetics, #MetLfe,#Fidelity and me) - #CMGnews
The Southern Computer Management Group (SCMG) held the very successful meeting with 40 attendees,
Main presentations (including mine) were uploaded to the site:
Main presentations (including mine) were uploaded to the site:
- AppDynamics: “The differing ways to monitor and instrument” LINK TO PRESENTATION
- DataKinetics: “Best Practices - Populating Big Data Repositories from DB2, IMS and VSAM” LINK TO PRESENTATION
- Cibra: “Capacity Management for Hybrid IT” LINK TO PRESENTATION
- Igor Trubin: "Is Your Capacity Available?” LINK TO PRESENTATION
- IBM: “IBM GTS Cirba Case Study “ LINK TO PRESENTATION
- MetLife: "Capacity Management Is Still Relevant” LINK TO PRESENTATION
- Fidelity: Too Big to Test: Breaking a production brokerage platform without causing financial devastation” LINK TO PRESENTATIION
Full Agenda is here:
Full Agenda is here:
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Friday, October 21, 2016
Interesting 1992 conference paper about SPC and machine learning written by Shewhart.
Interpreting statistical process control (SPC) charts using machine learning and expert system techniques
Conference Paper · June 1992
1st Mark Shewhart1.97 · Independent Researcher
Abstract
Statistical process control (SPC) charts are one of several tools
used in quality control. The SPC quality control tool has been
under-utilized due to the lack of experienced personnel able to identify
and interpret patterns within the control charts. The Special Projects
Office of the Center for Supportability and Technology Insertion (CSTI)
has developed a hybrid machine-learning and expert-system software tool
which automates the process of constructing and interpreting control
charts. The software tool draws control charts, identifies various chart
patterns, advises what each pattern means, and suggests possible
corrective actions. The application is easily modifiable for process
specific applications through simple modifications to the knowledge base
portion using any word processing software. The authors discuss control
charts, software functionality, software design, machine learning, and
the expert system
used in quality control. The SPC quality control tool has been
under-utilized due to the lack of experienced personnel able to identify
and interpret patterns within the control charts. The Special Projects
Office of the Center for Supportability and Technology Insertion (CSTI)
has developed a hybrid machine-learning and expert-system software tool
which automates the process of constructing and interpreting control
charts. The software tool draws control charts, identifies various chart
patterns, advises what each pattern means, and suggests possible
corrective actions. The application is easily modifiable for process
specific applications through simple modifications to the knowledge base
portion using any word processing software. The authors discuss control
charts, software functionality, software design, machine learning, and
the expert system
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Thursday, October 20, 2016
SCMG meeting in Cary, NC
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Monday, October 10, 2016
Interesting paper about "Adaptive Anomaly Detection in Cloud"
Adaptive Anomaly Detection in Cloud using Robust and Scalable Principal Component Analysis
by
Abstract
This paper proposes a novel and scalable model for automatic anomaly detection on a large system such as a cloud. Anomaly detection issues early warning of unusual behavior in dynamic environments by learning system characteristic from normal operational data. Anomaly detection in large systems is difficult to detect due heterogeneity, dynamicity, scalability, hidden complexity, and time limitation. To detect anomalous activity in the cloud, we need to monitor the datacenter and collect cloud performance data. In this paper, we propose an adaptive anomaly detection mechanism which investigates principal components of performance metrics. It transforms the performance metrics into a low-rank matrix and then calculates the orthogonal distance using the Robust PCA algorithm. The proposed model updates itself recursively learning and adjusting the new threshold value in order to minimize reconstruction errors. This paper also investigates the robust principal component analysis in distributed environments using Apache Spark as the underlying framework, specifically addressing cases in which a normal operation might exhibit multiple hidden modes. The accuracy and sensitivity of the model is tested on Google data center traces and Yahoo! datasets. The model achieves an 87.24% accuracy.
MY COMMENT: By the way the paper has referenced to MASF technique which I have enhanced and have been using (check my SETDS methodology) for years to capture anomalies (exceptions) and sudden short term trends against huge server farms (20,000+ servers) including private and public clouds. Note my way is much-much simpler and in spite the MASF has indeed a high rate of false positives, SETDS has the way to handle that well.
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Friday, October 7, 2016
Southern CMG meeting in Cary, NC on October 20th - Final Agenda
The SCMG is proud to announce our Fall 2016 Meeting, an all-day event on October 20, 2016
at
MetLife – Grace Hopper Auditorium (Bldg. 1 (MET 1), Floor 01, Room 600)
101 MetLife Way
Cary NC, 27513
at
MetLife – Grace Hopper Auditorium (Bldg. 1 (MET 1), Floor 01, Room 600)
101 MetLife Way
Cary NC, 27513
ð HURRY: REGISTER NO
LATER THAN OCT. 14!! This covers Breakfast and Lunch. We need to confirm the
number of Registrants so that we can properly plan the catering. We also
need your Registration information in order to get a list to MetLife Security
so that we will have visitor badges ready for you the morning of Oct. 20 when
you arrive in the MET1 lobby.
To REGISTER, use the registration page. You can use a PayPal account or
credit card. Your registration payment through the PayPal button
logs your registration.
AGENDA:
|
8:00-8:45 ET
|
Registration
/ Breakfast
|
Speaker BIO
|
|
8:45-9:30
|
Breakfast
provided with Sponsor Session: AppDynamics
|
|
|
9:30-10:30
|
Rick
Weaver
“Best
Practices - Populating Big Data Repositories from DB2, IMS and VSAM”
|
Over the
past 25 years, Rick Weaver has become a well-known mainframe expert
specializing in database protection, replication, recovery and performance.
Because of his vast expertise, he has authored numerous articles, whitepapers
and other valuable pieces on database technologies, and frequently spoken on
the subjects of database recovery and performance at conferences, symposiums
and user groups.
|
|
10:30-11:15
|
Ann
Dowling
“Capacity
Management Is Still Relevant”
|
Ann joined MetLife in April 2016 as the
Director of Capacity & Forecast Engineering. In this role Ann
will build on her extensive background working at IBM in various disciplines
including capacity planning, process architecture, performance engineering,
and offering management. Her professional passion is
Capacity Planning in support of the business and how it drives the
applications that consume resources on the IT infrastructure. Ann’s
most rewarding work has been leading teams to consolidate toward a common
‘best practices’ approach to capacity management. She did so for a
series of consolidations of independent data centers within IBM which evolved
into her role as the global Capacity Management process owner. That
work grounded Ann’s move to consulting services with external, non-outsourced
customers to evaluate their capacity management capabilities, identify
strengths and gaps to then build a roadmap for improvement. The
next step was working with a specific, large account to lead a team on the
implementation phases of the roadmap that gave Ann a more hands-on role
working directly with the engineering and operations teams and
management. She has been on the planning committee and speaker for
various IBM and CMG technical conferences. She was an instructor for
IBM’s Architecting for Performance class and author of a four-part series on
“Exploring Analytics to enable the Business and Service Value of Capacity
Planning”.
|
|
11:15-11:30
|
Break
|
|
|
11:30-12:30
|
Kyle
Parrish
CMG
2015 Mullen Award Winner
“Too
Big to Test: Breaking a production brokerage platform without causing
financial devastation”
|
Kyle currently works as a Director of
Technology Risk in the FI Information Security group at Fidelity
Investments. Kyle joined Fidelity in January of 2011 as a Director of
Performance Architecture charged with driving end-to-end testing of the
Fidelity Brokerage systems. Prior to joining Fidelity, Kyle worked as a
consultant for over 13 years, after a career in both the private sector and a
university research setting. Kyle’s roles have spanned everything from
program management to performance engineering to security, across industries
as varied as airlines, financial services, manufacturing, retail,
pharmaceuticals, and state government.
|
|
12:30-1:30
|
LUNCH
provided with Sponsor Session: Cirba
|
|
|
1:30-2:30
|
Igor Trubin
“Is Your Capacity Available?”
|
I
started my career in 1979 as an IBM/370 system engineer. In 1986 I got my
PhD. in Robotics at St. Petersburg Technical University (Russia) and then
worked as a professor teaching there CAD/CAM, Robotics and Computer Science
for about 12 years. I published 30 papers and made several presentations for
international conferences related to the Robotics, Artificial Intelligent and
Computer fields. In 1999 I moved to the US and worked at Capital One bank in
Richmond as a Capacity Planner. My first CMG paper was written and presented
in 2001. The next one, "Global and Application Level Exception Detection
System Based on MASF Technique," won a Best Paper award at CMG 2002 and
was presented again at UKCMG 2003 in Oxford, England. My CMG 2004 paper about
applying MASF technique to mainframe performance data was republished in the
IBM z/Series Expo. I also presented my papers in Central Europe CMG
conference and in numerous US regional meetings. I continue to enhance my
exception detection methodologies. After working more than 2 years as the
Capacity Management team lead for IBM, I had worked for SunTrust Bank for 3
years and then got back to IBM holding for 2+ years Sr. IT
Architect position. Currently I work for Capital One bank as IT Manager
for IT Capacity Management group. In 2015 I have been elected to the CMG (http://www.cmg.org) board
of directors. Blog: www.Trub.in
|
|
2:30-3:15
|
Shawn
Lundvall
“zBNA:
Theory and Overview”
|
I started my IBM career in 2001 in
Poughkeepsie in the Systems Architecture group writing the Principals of
Operations. In 2005 I got the opportunity to do hardware design of the fixed
point unit. In 2007 I moved to Richmond supporting clients as a Client Technical
Specialist. In 2013 I joined the Washington Systems Center as a Software
Engineer and am now a developer for zBNA and zPCR.
|
|
3:15-3:30
|
Break
|
|
|
3:30-4:30
|
Ken
Christiance
“IBM
GTS Cirba Case Study “
|
Ken Christiance –
Distinguished Engineer with 28 years’ experience in IBM. He has
been working in Strategic Outsourcing field since 1993; experience that spans
service management architectures, virtualization/server management and
analytics. Ken is currently a member of the Technology, Innovation and
Automation team that supports architecture and solution design for system
automation, virtualization and distributed server management. Ken is
patented and published for technologies that provide usage accounting and
billing, policy based automation, network design, virtualization and service
management tooling
|
|
4:30-5:30
|
SCMG
Committee Meeting
|
|
|
4:30-5:30
|
BOF
|
Optional: a breakout room will be reserved for folks who would like to
hold an impromptu BOF or post SCMG informal opportunity to network.
|
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Friday, September 9, 2016
41st International Performance & Capacity Conference
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Make an imPACt in 2016. Join us in La Jolla, CA this November
My presentation "Is your Capacity Available is scheduled to be presented at International CMG conference "imPACt 2016" in La Jolla, CA on Monday November 7th, 5-6pm. The abstract can be seen HERE
Please consider to attend!
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
My new presentation is scheduled at the Southern CMG meeting in Cary, NC on October 20th
I am glad to announce that our next SCMG meeting will be held on October 20th in Cary, NC.
See agenda and location details at the SCMG web page:
Tentative AGENDA:
Note I am presenting my new white paper there that is also scheduled to be published and presented at International CMG conference "imPACt 2016" in La Jolla, CA on Monday November 7th, 5-6pm. The abstract can be seen HERE
See agenda and location details at the SCMG web page:
Tentative AGENDA:
| 8:00-9:00 ET | Registration / Breakfast |
| 9:00-9:30 | Sponsor Session |
| 9:30-10:30 | Andrew Armstrong “tbd” |
| 10:30-11:30 | Ann Dowling “Capacity Management Is Still Relevant” |
| 11:30-12:30 | Kyle Parrish “Too Big to Test: Breaking a production brokerage platform without causing financial devastation”, CMG 2015 Mullen Award Winner |
| 12:30-1:30 | LUNCH with Sponsor Session: Cirba |
| 1:30-2:30 | Igor Trubin “Is Your Capacity Available?” |
| 2:30-3:30 | Shawn Lundvall “zBNA: Theory and Overview” |
| 3:30-4:30 | Ken Christiance “tbd” |
| 4:30-5:30 | SCMG Committee Meeting |
Please consider to attend both events!
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Wednesday, August 24, 2016
CMG #imPACt and #Velocity conferences cover the Capacity Management shift
UPDATE: I have invited Kevin to present again his famous Velocity presentation to newly formed NCACMG meet-up (will be on 12/6/2017 Wednesday in “IBM Washington Systems Center”, 470 Springpark
Place, Herndon, Va. 20170)
______________________________________
I am analyzing the content of the upcoming CMG conferences. ( imPACt 2017 )
Other related presentation was at last year CMG imPACt conference
Speaker: Ann Dowling
Company: MetLife
Session Title: A shift in who does capacity sizings
BTW I know very well both presenters. Follow this post for details!
______________________________________
I am analyzing the content of the upcoming CMG conferences. ( imPACt 2017 )
Most interesting was that the upcoming Velocity conference also covers this subject:
Speaker: Kevin McLaughlin
Company: Capital One
Session Title: Is capacity management still needed in the public cloud?Session Abstract:
The cloud holds the promise of bottomless capacity, available instantly. Recently, Capital One has been shifting a significant portion of its workload to the public cloud. Kevin McLaughlin explores what capacity management looks like in the cloud, which old concepts still apply, which should be retired, and what new metrics become important and covers the importance of performance management. Kevin also outlines what needs to be monitored as workloads transition to the cloud and what to monitor once a workload is fully in the cloud, as well as considerations for ensuring the legacy environment maintains sufficient capacity during the transition.Other related presentation was at last year CMG imPACt conference
Speaker: Ann Dowling
Company: MetLife
Session Title: A shift in who does capacity sizings
Session Abstract:
The accelerating advance of hybrid infrastructures with promotion of self-service requests for capacity is causing a significant shift in who is responsible for sizing capacity requirements. The shift is moving out of the direct management by IT infrastructure capacity planners out to the end user or consumer - often application owners and development teams. This can be viewed as a positive shift that activates the linkage between infrastructure and application teams. The challenge is to ensure the requestors have the skills and tools to adequately size their requirements for cpu, memory, and storage for both their immediate needs and with an understanding of workload growth patterns along with the financial implications. This presentation will focus on the organizational and cultural shifts that result from the emergence and popularity of self-service capacity sizings and requests.
BTW I know very well both presenters. Follow this post for details!
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Friday, August 5, 2016
My paper "Is your Capacity Available?" has been accepted for the imPACt 2016 by CMG Conference to be held November 7 - 10, 2016 in La Jolla CA.
You may check the abstract and some additional information about the paper here: https://www.researchgate.net/publication/281389752_Is_Your_Capacity_Available or in this blog: http://www.trub.in/2013/05/is-your-capacity-available-topic-for.html (#CMGimPACt #capacity)
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Monday, August 1, 2016
CONFERENCE ON AVAILABILITY, RELIABILITY AND SECURITY
The 11th International Conference on Availability, Reliability and Security (“ARES”) will bring together researchers and practitioners in the area of dependability. ARES will highlight the various aspects of security - with special focus on the crucial linkage between availability, reliability and security.
Interesting..., but I do not see any topics about the availability and capacity interconnection in this forum...
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Thursday, July 28, 2016
Adrian Cockcroft's Blog: My CMG paper on Crunching Data In the Cloud is pub...
Adrian Cockcroft's Blog: My CMG paper on Crunching Data In the Cloud is pub...: The slides are also available at http://www.slideshare.net/adrianco/crunch-your-data-in-the-cloud-with-elastic-map-reduce-amazon-emr-hadoop ...
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Friday, July 8, 2016
Performance Problem Diagnosis in Cloud Infrastructures
I have been notified by RG about new reference to my paper "Capturing Workload Pathology by Statistical Exception Detection System". The following interesting tithes referenced my work:
by
Umeå University
Abstract
Cloud datacenters comprise hundreds or thousands of disparate application services, each having stringent performance and availability requirements, sharing a finite set of heterogeneous hardware and software resources. The implication of such complex environment is that the occurrence of performance problems, such as slow application response and unplanned downtimes, has become a norm rather than exception resulting in decreased revenue, damaged reputation, and huge human-effort in diagnosis. Though causes can be as varied as application issues (e.g. bugs), machine-level failures (e.g. faulty server), and operator errors (e.g. mis-configurations), recent studies have attributed capacity-related issues, such as resource shortage and contention, as the cause of most performance problems on the Internet today. As cloud datacenters become increasingly autonomous there is need for automated performance diagnosis systems that can adapt their operation to reflect the changing workload and topology in the infrastructure. In particular, such systems should be able to detect anomalous performance events, uncover manifestations of capacity bottlenecks, localize actual root-cause(s), and possibly suggest or actuate corrections.
This thesis investigates approaches for diagnosing performance problems in cloud infrastructures. We present the outcome of an extensive survey of existing research contributions addressing performance diagnosis in diverse systems domains. We also present models and algorithms for detecting anomalies in real-time application performance and identification of anomalous datacenter resources based on operational metrics and spatial dependency across datacenter components. Empirical evaluations of our approaches shows how they can be used to improve end-user experience, service assurance and support root-cause analysis.
This thesis investigates approaches for diagnosing performance problems in cloud infrastructures. We present the outcome of an extensive survey of existing research contributions addressing performance diagnosis in diverse systems domains. We also present models and algorithms for detecting anomalies in real-time application performance and identification of anomalous datacenter resources based on operational metrics and spatial dependency across datacenter components. Empirical evaluations of our approaches shows how they can be used to improve end-user experience, service assurance and support root-cause analysis.
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga










