My old idea of building/using IT-Control chart as a good weekly pattern visualization of users' behavior now is seen in Facebook!
This blog relates to experiences in the Systems Capacity and Availability areas, focusing on statistical filtering and pattern recognition and BI analysis and reporting techniques (SPC, APC, MASF, 6-SIGMA, SEDS/SETDS and other)
Popular Post
-
I have got the comment on my previous post “ BIRT based Control Chart “ with questions about how actually in BIRT the data are prepared for ...
-
Re-posting interesting article from R-vs.-Python-for-Data-Science R vs. Python for Data Science Norm Matloff, Prof. of Computer ...
_
Sunday, January 10, 2021
My Weekly IT-Control chart finally is used by Facebook
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Monday, December 28, 2020
My Article: "IT-Control Chart" reached 800 reads (#qualitycontrol #capacitymanagement #performanceengineering)
ABSTRACT: The Control Chart is one of the main Six Sigma tools to optimize business processes. After some adjustments it is used now as visualization tool in IT Capacity Management especially in “behavior learning” products to underline performance and capacity usage anomalies. This review answers the following questions. What is the Control Chart and how to read it and where to use? Review of some performance tools that use it. Control chart types: MASF charts vs. classical SPC; introduction to IT-Control Chart for IT application performance control. How to build a Control Chart using Excel for interactive analysis and R scripting to do it automatically?
The full paper is HERE
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Wednesday, December 9, 2020
#AWS #Re:invent - Capacity Planning Use Case
From the following session:
How Capital One manages the health of its applications on AWS
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Friday, December 4, 2020
This year at #CMGIMPACT2021, I’ll be a speaker on "Cloud Resources Workload Profiling". Anyone looking to attend please contact me for a discount code, saving 40% off the registration price. #CMGnews
This year at #CMGIMPACT2021, I’ll be a speaker on Cloud Resources Workload Profiling. Anyone looking to attend please contact me for a discount code, saving 40% off the registration price. Find more information on IMPACT 2021 Virtual Conference here: cmgimpact.com/home2021/
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
#CMGnews: "Detection of Performance #Anomaly using DESOM" - #cmgimpact2021 session
Deep Embedded Self Organizing Map (DESOM), a hybrid Deep Neural Network based Autoencoder-Decoder (AE-DE) with an embedded Self Organizing Map (SOM), is applied successfully for the first time to detect anomaly in the performance metrics of mobile network entities with over 94% accuracy. SOM has been widely used in many areas for anomaly detection such as fraud detection, intrusion detection, etc. DESOM is a recent enhancement of SOM but not evaluated as practical solution for real problems prior to this work. Several novel methods to detect concept drift using the intrinsic features of DESOM have been incorporated in the complete solution pipeline.
Speaker
Jayanta Choudhury
Senior Data Scientist
Ericsson Inc.
Santa Clara, California United States
Anila Joshi
Sr. Data Science Manager
Ericsson Inc.
Santa Clara, California United States
Track
Performance Engineering and DevOps
https://cmgimpact.com/detection-of-performance-anomaly-using-desom/
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Tuesday, December 1, 2020
#AWS #Re:invent virtual conference activities log
Machine Learning Keynote (https://virtual.awsevents.com/media/1_07cg4srl)
Productionizing R workloads using Amazon SageMaker, featuring Siemens
"It’s easier than ever to grow your compute capacity and enable new types of cloud computing applications while maintaining the lowest total cost of ownership (TCO) by blending EC2 Spot Instances, On-Demand Instances, and Savings Plans purchase models. In this session, learn how to use the power of EC2 Fleet with AWS services such as Amazon EC2 Auto Scaling, Amazon ECS, Amazon EKS, Amazon EMR, and AWS Batch to programmatically optimize costs while maintaining high performance and availability. Dive deep into cost-optimization patterns for workloads such as containers, web services, CI/CD, batch, big data, and more."
Most interesting slides from that session:
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
#AWS #Re:invent virtual conference activities (I support Capital One's sponsor booth there! )
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Friday, November 6, 2020
"Cloud Resources Workload Profiling" - my new presentation #cmgimpact2021 (#CMGnews #CloudComputing )
Abstract: How to be sure a cloud object’s (e.g, AWS EC2, RDS or EBS) workload fits the rightsized resources (Compute, RAM, IO/s and Network traffic)? It is very difficult to do using raw system performance data from monitoring tools. The best way to do that is using a weekly workload profile, which is a graphical visualization in form of MASF IT-Control chart. This chart shows the stability of the workload, reveals the anomalies happened recently, such as run-away, memory leaks or specifically important for cloud objects, the unusual number of hours the object is down all compared with the usual weekly pattern.
This presentation will describe how to build, read, and use workload profiles using real data examples and demonstrates how cloud capacity scaling could be verified.
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Wednesday, October 7, 2020
"Optimization & Improving Performance to Keep Clouds Light" - another capitalone.com/tech/ post with my contribution
A new article has just been issued with a reference to my previous capitalone.com/tech/ post (see
My External Post on Capitalone.com/tech/ - Optimizing Your Public Cloud for Maximum Efficiency
) and with my short quote:
"...Data Analyst Igor Trubin, explains the mechanics behind data collection. “We found a way to collect adequate cloud usage performance data to automatically recognize the workload patterns of different cloud objects and their subsystems (clusters of servers, databases, containers, disk volumes and networks), including the cost of using them.” Some details of the method for doing that we shared in the following tech blog post: Optimizing Your Public Cloud for Maximum Efficiency..."
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Wednesday, September 16, 2020
To Work from Home or to Work from Work?
Below are my comments to the following post of my former manager:
What do I miss? (working 100% at home)
Whiteboarding. Definitely ! I used to do
that a lot:
- explaining my ideas to developers, so
they could implement what want not blindly but with passion. Now working from home I
spent twice more time via zoom and still not sure I ignite that passion;
- proving my new and innovative
concepts to my boss. Now they have to listen or reading my Ruglish,
respectively my ability to convince accepting my ideas is declining...
Serendipit ....Less impactful hallway
or over-the-cube-wall conversations can also solve problems and create bonds.....
I feel that too. But in
DevOps-Adgile-Slack-zoom environment I am getting used to get what I need. My
hobby to be always on-line blogging-presenting (Now it is Virtual
Convergence/Seminar) helps a lot.
Rhythms. Yes, life circles are changed. But maybe
it is for good. At least for me it is not bad to break current routine and
establish new ones. With kids growing up it is anyway
unavoidable....
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Tuesday, September 15, 2020
My CMG Video presentation "Catching Anomaly and Normality in Cloud by Neural Net and Entropy Calculation"
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Saturday, August 22, 2020
CPD - Change Point Detection (#ChangeDetection) is implemented in the free web tool Perfomalist
UPDATE 11/21/2021
The method is implemented as a Perfomalist API: https://www.trutechdev.com/2021/11/the-change-points-detection-perfomalapi.html.
Note there a tuning parameters that corresponds to once explained below:
- sValue - Statistical band in %, where 100 is UCL=MAX, 0 is UCL=LCL=mean). - N - normality confidence band;
- eValue - Exception Value (EV) threshold in % of actual historical average. - I - model insensitivity;
- BaseLineLength - The time period to compare current value against.
_______________________________________
The next version of the Perfomalist (https://www.perfomalist.com/ ) is coming and will include a new functionality - Change Point Detection.
How to find a change in the historical time-series data?
Long ago I have developed a method to do that which is based on EV data (Exception Value - a magnitude of anomalies collected historically).
Idea: any change that occurred first would appear as an anomaly and then become a normality (norm), so collecting and analyzing the severity of all anomalies opens the possibility to find phases in the history with different patterns. To detect that mathematically one just needs to find all roots of the following equation: EV(t)=0 , where t is time. But it is too simple as that might give you too many change points. To control the the sensitivity of detecting change points the method should have some sensitivity tuning parameters, such as following:
N - normality confidence band in percentiles = UCL-LCL (if it is 100%, that means all observations is normal, 0% means all observations abnormal)
|EV(t,N)|=I or
|EV(t,UCL-LCL)|=I
Where UCL is upper control limit and LCL is lower control limit. Why "||" (absolute value)? To catch two types of changes: going up- and downwards.
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga




