The presentation will be published in SCMG site: http://regions.cmg.org/regions/scmg/fall_09/richmond/meeting_09_24_09b.htm
This blog relates to experiences in the Systems Capacity and Availability areas, focusing on statistical filtering and pattern recognition and BI analysis and reporting techniques (SPC, APC, MASF, 6-SIGMA, SEDS/SETDS and other)
Popular Post
-
I have got the comment on my previous post “ BIRT based Control Chart “ with questions about how actually in BIRT the data are prepared for ...
-
Re-posting interesting article from R-vs.-Python-for-Data-Science R vs. Python for Data Science Norm Matloff, Prof. of Computer ...
_
Sunday, September 20, 2009
Near-Real-Time IT-Control Charts
On the next Thursday September 24, 2009 in the Richmond's SCMG meeting I am going to present my updated version of previous presentation called "Power of Control Chart". This time the focus is on Near-Real-Time IT-Control Charts. Below is the clip that shows the example of Near-Real-Time IT-Control Chart simulated by R-program:
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Thursday, August 27, 2009
IT-Chart: The Best Way to Visualize IT Systems Performance
How to see most current metric data, most recent one and also in retrospective, but in one picture? Is that possible? Yes, it is.
I guess the simple radar in a plane or ship cockpit refreshes current data on a top of most recent and shows approaching "future". The SEDS control chart is similar and uses the border line to separate current data and most recent one. Plus it gives a historical base-line to show you what can be expected and for comparison.
I believe it is powerful way to visualize IT Systems Performance, so I made up a name for that chart: "IT-CHART".
(Not only because IT is my initials....)
I plan to add the pictures from this blog as additional slides to my CMG'09 workshop, which includes the R-script to build IT-CHART based on CSV input. (see abstract here)
I guess the simple radar in a plane or ship cockpit refreshes current data on a top of most recent and shows approaching "future". The SEDS control chart is similar and uses the border line to separate current data and most recent one. Plus it gives a historical base-line to show you what can be expected and for comparison.I believe it is powerful way to visualize IT Systems Performance, so I made up a name for that chart: "IT-CHART".
(Not only because IT is my initials....)
I plan to add the pictures from this blog as additional slides to my CMG'09 workshop, which includes the R-script to build IT-CHART based on CSV input. (see abstract here)
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Thursday, July 23, 2009
Real-Time Control Charts for SEDS
I still analyze different tools that capture computer application abnormalities based on real-time data. In addition to Integrien (now it is a part of VMware's tool called AliveVM ) and Netuitive I have recently looked at BMC ProActive Net Analytics. I have spoken with BMC SMEs and they showed me a live demo of the tool. I always respected BMC (and espetialy BGS) as actually the inventor of this approach (MASF) and long ago I used to analyze statistical exceptions using BMC Visualizer and BMC Perceive (BTW I have published in my papers a few examples how I did that) . Now they have another and very good tool for the same purpose (http://documents.bmc.com/products/documents/49/13/84913/84913.pdf)Watching the live presentation I got a positive impression of how that works for complex applications and transactions correlating different abnormal events with possibility to reduce false positives situations. Interesting that the combination of dynamic and static thresholds are used there to generate alarms. Just like SEDS does - static one to capture hot issues (run-aways and leaks) and statistical ones for early warnings.
Now I have a very difficult task to choose from those three products (plus SEDS) to recommend to my management...
Speaking about SEDS, I have decided to play with near-real time data to see how difficult would be to redesign SEDS making it works more similar with mentioned above modern and serious tools. Fortunately SEDS is just a bunch of SAS Marcos with parameters which helped me to make the adjustment needed to include today's data. And surprisingly that was pretty easy task! I spent only a couple days to developed a "real-time SEDS" prototype. Currently what it only does is building every hour the real-time Control Charts that can be seen at the beginning of this post.
I plan to include some details about real-time Control Charts to my upcoming CMG'09 Workshop.
Labels:
Control Chart,
Integrien,
ProActive Net,
Real-time monitoring,
SEDS
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Thursday, July 16, 2009
The Performance and Capacity Analyst Bookshelf
0. PERSONAL: Alex Podelko's Capacity/Performance links collection. - Richest in the Internet!
1. ON-LINE
- Guerrilla Capacity Planning by Neil J. Gunther, M.Sc., Ph.D. and Guerrilla Capacity Planning PART II: Weapons of Mass Instruction by Neil J. Gunther
- Ray Wicks: Getting Started in z/OS Capacity Planning
Part2 Getting Started in z/OS Capacity Planning
Part3 Getting Started in z/OS Capacity Planning
Part4 Getting Started in z/OS Capacity Planning
Part5 Getting Started in z/OS Capacity Planning
2. TO BUY
- John Allspaw:The Art of Capacity Planning ...
- The Performance and Capacity Analyst Bookshelf by Rick Ralston and Dan Schwarz
1. ON-LINE
- Guerrilla Capacity Planning by Neil J. Gunther, M.Sc., Ph.D. and Guerrilla Capacity Planning PART II: Weapons of Mass Instruction by Neil J. Gunther
- Ray Wicks: Getting Started in z/OS Capacity Planning
Part2 Getting Started in z/OS Capacity Planning
Part3 Getting Started in z/OS Capacity Planning
Part4 Getting Started in z/OS Capacity Planning
Part5 Getting Started in z/OS Capacity Planning
2. TO BUY
- John Allspaw:The Art of Capacity Planning ...
- The Performance and Capacity Analyst Bookshelf by Rick Ralston and Dan Schwarz
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Monday, July 6, 2009
My CMG’09 Sunday Workshop
My workshop entitled
"Power of Control Charts: How to Read, How to Build, How to Use”
has been accepted for the CMG’09 Sunday Workshop program to be held at the Gaylord Texan in Dallas, Texas, December 6, 2009 (http://cmg.org/conference/cmg2009/)
The workshop proposal is following:
One of the most powerful ways to visualize computer system behavior is the Control Chart. Originally used in Mechanical Engineering, it has become one of the main Six Sigma tools to optimize business processes, and after some adjustments it is used in IT Capacity Management area especially in “behavior learning” products.
During the workshop the following topics will be discussed: What is the Control Chart? Where the Control Chart is used: review of some systems performance tools that use it. Control chart types: MASF charts vs. SPC. Gallery of already published charts in CMG papers plus some new charts with explanations on how to read them. How to build a Control Chart: using Excel for interactive analysis and R to do it automatically. The session includes a live demonstration of Excel to build different types of control charts against real performance data. Attendees will be provided CDs with the data in spreadsheets and will build Control Charts themselves even with their own data. Finally, they will be able to run an R-script to build a Control Chart based on input CSV data.
This workshop is based on series of CMG papers published by the author. The prototype of the workshop was presented twice this year in Southern CMG meetings in VA and NC.
The presentation slides are already published here: https://www.researchgate.net/profile/Igor_Trubin/publication/259486489_TrubinCMG2009_IT_ControlCharts_SCMG_Fall/data/59c14e9e0f7e9b21a82657b6/CMG2009-Workshop-Trubin.pptx
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Thursday, July 2, 2009
Capacity Management Found in Translation
I have just created my 2nd blog to share my technical ideas and thoughts in Russian. If you can read Russian please visit http://www.ukor.blogspot.com/.
The name of my new blog is "Управление Вычислительной Мощностью" which simply means “Capacity Management”. That term translation I have recently found in a Russian article (click here to read) which was published in 2008 by Enterprise Systems and Software Laboratory, HP Laboratories Palo Alto. I was so glad that I had finally figured out how “Capacity Management” is said in Russian! The past 10 years doing Capacity Management I always had a problem explaining to my Russian friends and relatives what my occupation was! Now I know and that fact inspired me to start my new blog for Russian readers.
Another reason is the 20th anniversary of my 1st program which I wrote and sold. That was the graphical editor with some CAD features I wrote using FORTRAN for PC with PDP type of processor (DVK-3). The name of that program was UKOR (In Russian that means “REPROOF”). That’s why the link to my new blog is “ http://www.UKOR.blogspot.com/ “!
The name of my new blog is "Управление Вычислительной Мощностью" which simply means “Capacity Management”. That term translation I have recently found in a Russian article (click here to read) which was published in 2008 by Enterprise Systems and Software Laboratory, HP Laboratories Palo Alto. I was so glad that I had finally figured out how “Capacity Management” is said in Russian! The past 10 years doing Capacity Management I always had a problem explaining to my Russian friends and relatives what my occupation was! Now I know and that fact inspired me to start my new blog for Russian readers.
Another reason is the 20th anniversary of my 1st program which I wrote and sold. That was the graphical editor with some CAD features I wrote using FORTRAN for PC with PDP type of processor (DVK-3). The name of that program was UKOR (In Russian that means “REPROOF”). That’s why the link to my new blog is “ http://www.UKOR.blogspot.com/ “!
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Wednesday, June 17, 2009
Management by Exception: Business vs. System
Management by Exception is actually an old idea and it is used for the Business Process management and even for the Accounting as defined in the following website that I have recently found: http://www.allbusiness.com/glossaries/management-by-exception/4944378-1.html
Wikipedia, referring to the same source, defines Management by Exception as a
"policy by which management devotes its time to investigating only those situations in which actual results differ significantly from planned results. The idea is that management should spend its valuable time concentrating on the more important items (such as shaping the company's future strategic course). Attention is given only to material deviations requiring investigation."
I would say if one applies this definition to IT, it turns to my term "System Management by Exception" where the "management" is Capacity management analysts or Capacity planners and "material deviations" are servers or applications' exceptions.
Speaking about applications' exceptions, currently I am working on applying “Management by Exception” approach not to servers farm capacity management (I think I have already done this successfully) but to a set of applications to produce automatically the list of only those applications that are having some exceptions (unusual but not yet deadly behavior) to help providing proactive application capacity/performance management. Why? Because in some IT environments with large number of applications the centralized capacity management does not exist and application support teams have to play that role and SEDS (System Management by Exception tool) should deliver automatically them what systems need attention within each exceptional application.
Wikipedia, referring to the same source, defines Management by Exception as a
"policy by which management devotes its time to investigating only those situations in which actual results differ significantly from planned results. The idea is that management should spend its valuable time concentrating on the more important items (such as shaping the company's future strategic course). Attention is given only to material deviations requiring investigation."
I would say if one applies this definition to IT, it turns to my term "System Management by Exception" where the "management" is Capacity management analysts or Capacity planners and "material deviations" are servers or applications' exceptions.
Speaking about applications' exceptions, currently I am working on applying “Management by Exception” approach not to servers farm capacity management (I think I have already done this successfully) but to a set of applications to produce automatically the list of only those applications that are having some exceptions (unusual but not yet deadly behavior) to help providing proactive application capacity/performance management. Why? Because in some IT environments with large number of applications the centralized capacity management does not exist and application support teams have to play that role and SEDS (System Management by Exception tool) should deliver automatically them what systems need attention within each exceptional application.
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Wednesday, June 10, 2009
CMG Board of Directors Nomination
2020 UPDATE:
Final update:
Thanks for all who voted for me last year! I am resubmitting my nomination again for this year.
2014 UPDATE:
This year I was nominated again !
So, if you are a CMG.org member, please vote! How to vote check HERE.
By Chair of 2009 CMG Nominating Committee I have been asked to nominate myself to CMG Board of Directors. Apparently I am qualified for that and I believe it is a great honor. I have decided to do that and below it is my nomination statement.
Willingness to Serve:
CMG has been an extremely valuable part of my professional life for the past ten years. Because of CMG, I became a known specialist in IT Capacity Management discipline! I have already worked at the local level to support the organization and would like to serve on CMG's Board of Directors to continue promoting the organization throughout the IT community. My company and family members support my involvement with and commitment to CMG.
Professional Work Experience:I have over 30 years of experience in the IT field. I have started my career in 1979 as an IBM 370 system engineer. In 1986, I received my PhD. in Robotics at St. Petersburg Technical University (Russia), where I then taught full-time such subjects as CAD/CAM, Robotics and Computer Science for about 12 years. I have published more than 30 papers and made several presentations for different international conferences related to the Robotics, Artificial Intelligence and Computer fields. In 1999, I moved to the US and worked at Capital One bank as a Capacity Planner. My first CMG paper was written and presented in 2001. The next one, "Global and Application Level Exception Detection System Based on MASF Technique," won a Best Paper award at CMG’02 and was presented again at UKCMG’03 in Oxford, England. My CMG’04 was republished in the IBM z/Series Expo. I also presented my papers in Central Europe CMG conference (Austria) and at numerous US regional meetings. After working more than two years as the Capacity Management Team Lead for IBM, in 2007 I have accepted a Senior Capacity Planner position at SunTrust Bank where I am currently employed.
Other Professional Experience:I have a long experience working as a programmer. I have also acquired extensive managerial experience working as the Head of the CAD/CAM University’s lab and Team Lead at IBM. Since March 2005, I have been severing as Vice Chair of Southern CMG, providing vendors connections.
Candidate Statement:I believe that I am uniquely qualified and motivated to serve CMG and its future development, as the IT landscape changes. My major accomplishment is Statistical Exception Detection System (SEDS) for IT Capacity Management. SEDS ideas and techniques are published in a series of my CMG papers over the period of the last ten years and also in this technical blog. My position as a Capacity Management expert and my dedication to the CMG organization will allow me to contribute in substantial ways. I further believe that my teaching experience could enhance CMG’s training and educational services for technical community. If elected, I will diligently pursue innovative ways to strengthen the organization’s membership. I will continue the CMG’s dedicated tradition of volunteerism and will actively seek ways to support and improve CMG's commitment to supporting its members.
If you are CMG member, please vote for me!
#CMGnews: I have been re-elected again to Computer Measurement Group (www.CMG.org) #BoardOfDirectors
Final update:
The COMPUTER MEASUREMENT GROUP (www.CMG.org) membership has elected me to serve as Director for the 2016 - 2017 term
2015 UPDATE:Thanks for all who voted for me last year! I am resubmitting my nomination again for this year.
2014 UPDATE:
This year I was nominated again !
So, if you are a CMG.org member, please vote! How to vote check HERE.
By Chair of 2009 CMG Nominating Committee I have been asked to nominate myself to CMG Board of Directors. Apparently I am qualified for that and I believe it is a great honor. I have decided to do that and below it is my nomination statement.
Willingness to Serve:
CMG has been an extremely valuable part of my professional life for the past ten years. Because of CMG, I became a known specialist in IT Capacity Management discipline! I have already worked at the local level to support the organization and would like to serve on CMG's Board of Directors to continue promoting the organization throughout the IT community. My company and family members support my involvement with and commitment to CMG.
Professional Work Experience:I have over 30 years of experience in the IT field. I have started my career in 1979 as an IBM 370 system engineer. In 1986, I received my PhD. in Robotics at St. Petersburg Technical University (Russia), where I then taught full-time such subjects as CAD/CAM, Robotics and Computer Science for about 12 years. I have published more than 30 papers and made several presentations for different international conferences related to the Robotics, Artificial Intelligence and Computer fields. In 1999, I moved to the US and worked at Capital One bank as a Capacity Planner. My first CMG paper was written and presented in 2001. The next one, "Global and Application Level Exception Detection System Based on MASF Technique," won a Best Paper award at CMG’02 and was presented again at UKCMG’03 in Oxford, England. My CMG’04 was republished in the IBM z/Series Expo. I also presented my papers in Central Europe CMG conference (Austria) and at numerous US regional meetings. After working more than two years as the Capacity Management Team Lead for IBM, in 2007 I have accepted a Senior Capacity Planner position at SunTrust Bank where I am currently employed.
Other Professional Experience:I have a long experience working as a programmer. I have also acquired extensive managerial experience working as the Head of the CAD/CAM University’s lab and Team Lead at IBM. Since March 2005, I have been severing as Vice Chair of Southern CMG, providing vendors connections.
Candidate Statement:I believe that I am uniquely qualified and motivated to serve CMG and its future development, as the IT landscape changes. My major accomplishment is Statistical Exception Detection System (SEDS) for IT Capacity Management. SEDS ideas and techniques are published in a series of my CMG papers over the period of the last ten years and also in this technical blog. My position as a Capacity Management expert and my dedication to the CMG organization will allow me to contribute in substantial ways. I further believe that my teaching experience could enhance CMG’s training and educational services for technical community. If elected, I will diligently pursue innovative ways to strengthen the organization’s membership. I will continue the CMG’s dedicated tradition of volunteerism and will actively seek ways to support and improve CMG's commitment to supporting its members.
If you are CMG member, please vote for me!
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Thursday, May 7, 2009
SEDS charts at SCMG
SCMG has just held two great meetings:
My presentation "Power of Control Charts" was well received. The slides are published there: https://www.researchgate.net/profile/Igor_Trubin/publication/259486489_TrubinCMG2009_IT_ControlCharts_SCMG_Fall/data/59c14e9e0f7e9b21a82657b6/CMG2009-Workshop-Trubin.pptx
I was able to demonstrate in live some control charts building technique including R scripting. It encouraged me to submit workshop proposal for CMG'09 "Power of Control Charts: How to Read, How to Build, How to Use".

If it's accepted, please come to see my workshop in December 6 in Dallas, Texas : http://www.cmg.org/conference/!
P.S. Interesting news I got from SCMG meeting about R: The SAS v. 9.2 can execute R scripts. Does anybody try?
P.S. Interesting news I got from SCMG meeting about R: The SAS v. 9.2 can execute R scripts. Does anybody try?
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Wednesday, March 25, 2009
Performance Anomaly ("Perfomaly") Detection. Parts 1-4: Power of Control Charts
_______
Based on my old workshop (Power of Control Chart), which I ran a few times a several years ago I develop the updated version of it and that will be the part of a training course
Performance Anomalies ("Perfomalies") Detection.
That will consist of the following parts:
So the parts 1-4 is the updated version of my old workshop about:
Based on my old workshop (Power of Control Chart), which I ran a few times a several years ago I develop the updated version of it and that will be the part of a training course
Performance Anomalies ("Perfomalies") Detection.
That will consist of the following parts:
1. Introduction to Performance
Anomaly Detection
2. Detecting performance
anomalies by Control Charts - lecture.
3. Building control charts by
using Excel - hands-on exercises.
4. Detecting performance
anomalies by Control charts using R on cloud server ( AWS ) - hands-on
exercises.
- includes the Instruction video how to build R environment on AWS cloud
5. Detecting Novelties in
performance data by using Exception Value (EV) approach (type of “knee”
detection) - lecture
6. Detecting Novelties in
performance data - hands-on exercises.
7. Detecting normality in the
performance workload data by neural nets and deep learning – lecture’
8. Detecting normality by using R
and R NN packages - hands on exercises.
9. Detecting anomalous short
living objects by using entropy calculation - lecture
10. Detecting anomalous short
living objects – hands-on exercises.
So the parts 1-4 is the updated version of my old workshop about:
- What is the Control Chart? - A little bit of theory and history.
- Where the Control Chart is used: Review of some systems performance tools on a market that built and use control charts.
- How SEDS uses that - MASF charts vs. SPC ones; long gallery of already published charts in the CMG papers plus some new ones with explanations how to read them.
- How to build Control chart: using Excel for interactive analysis and R to automate the control charts generating with live demonstration of the technique.
- NEW: How to build the R environment (Rstudio) in the cloud (AWS) server (EC2) and using R code to build control charts against your own data.
Data and R script for testing AWS EC2 with R environment and for the 1st R exercise:
Below is the data in CSV format (supposed to copy to test.csv file) and simple R script to build the monthly profile of some real Unix file system space utilization in form of a monthly Control Chart.
- How SEDS uses that - MASF charts vs. SPC ones; long gallery of already published charts in the CMG papers plus some new ones with explanations how to read them.
- How to build Control chart: using Excel for interactive analysis and R to automate the control charts generating with live demonstration of the technique.
- NEW: How to build the R environment (Rstudio) in the cloud (AWS) server (EC2) and using R code to build control charts against your own data.
Data and R script for testing AWS EC2 with R environment and for the 1st R exercise:
Below is the data in CSV format (supposed to copy to test.csv file) and simple R script to build the monthly profile of some real Unix file system space utilization in form of a monthly Control Chart.
TEST.SCV
day,CurrentMonthData,UpLimit,Mean,LowLimit
1,0.45,0.54,0.42,0.31
2,0.45,0.54,0.42,0.31
3,0.45,0.54,0.42,0.31
4,0.45,0.54,0.42,0.31
5,0.45,0.54,0.42,0.31
6,0.45,0.53,0.43,0.32
7,0.45,0.54,0.43,0.32
8,0.45,0.54,0.43,0.32
9,0.45,0.53,0.43,0.33
10,0.45,0.53,0.43,0.33
11,0.45,0.53,0.43,0.33
12,0.72,0.53,0.43,0.33
13,0.72,0.53,0.43,0.33
14,0.72,0.53,0.42,0.32
15,0.45,0.53,0.42,0.32
16,0.45,0.55,0.43,0.31
17,0.45,0.55,0.44,0.33
18,1.00,0.54,0.44,0.33
19,0.84,0.54,0.44,0.33
20,0.84,0.54,0.44,0.34
21,0.84,0.54,0.44,0.34
22,,0.54,0.44,0.34
23,,0.52,0.44,0.36
24,,0.52,0.44,0.36
25,,0.51,0.43,0.36
26,,0.66,0.46,0.26
27,,0.66,0.46,0.25
28,,0.62,0.45,0.28
29,,0.62,0.45,0.28
30,,0.54,0.43,0.32
31,,0.54,0.43,0.32
## R script to plot control chart CSV input - I.Trubin
###############################################################
cchrt=read.table('test.csv', header=T, sep=",")
## R script to plot control chart CSV input - I.Trubin
###############################################################
cchrt=read.table('test.csv', header=T, sep=",")
plot( cchrt[,1],cchrt[,2], type="l",col="black", ylim=c(0,1),lwd=2,ann=F)
points(cchrt[,1],cchrt[,3],type="l",col="red", ylim=c(0,1),lwd=1,ann=F)
points(cchrt[,1],cchrt[,4],type="l",col="green", ylim=c(0,1),lwd=1,ann=F)
points(cchrt[,1],cchrt[,5],type="l",col="blue", ylim=c(0,1),lwd=1,ann=F)
mtext("# of transactions (K)", side=2, line=3.0)
mtext("days of month", side=1, line=3.0)
mtext("CONTROL CHART", side=3, line=1.0)
legend(9,0.3,c("Current Month","UpperLimit","Mean","LowerLimit"),
col=c("black","red","green","blue"),lwd=c(2,1,1,1),bty="n")
###############################################################
points(cchrt[,1],cchrt[,4],type="l",col="green", ylim=c(0,1),lwd=1,ann=F)
points(cchrt[,1],cchrt[,5],type="l",col="blue", ylim=c(0,1),lwd=1,ann=F)
mtext("# of transactions (K)", side=2, line=3.0)
mtext("days of month", side=1, line=3.0)
mtext("CONTROL CHART", side=3, line=1.0)
legend(9,0.3,c("Current Month","UpperLimit","Mean","LowerLimit"),
col=c("black","red","green","blue"),lwd=c(2,1,1,1),bty="n")
###############################################################
Result is in the picture.
(Other examples posted here: Near-Real-Time IT-Control Charts )
If you would like to attend my workshop - put your contact information to the comment of the post.
Raw time-series CSV data for the case is below:
MonthlyRawData.csv
date,metric
6/1/2008,0.39
6/2/2008,0.39
6/3/2008,0.39
6/4/2008,0.39
6/5/2008,0.39
6/6/2008,0.39
6/7/2008,0.39
6/8/2008,0.39
6/9/2008,0.39
6/10/2008,0.39
6/11/2008,0.39
6/12/2008,0.39
6/13/2008,0.39
6/14/2008,0.39
6/15/2008,0.39
6/16/2008,0.39
6/17/2008,0.39
6/18/2008,0.39
6/19/2008,0.39
6/20/2008,0.39
6/21/2008,0.39
6/22/2008,0.39
6/23/2008,0.4
6/24/2008,0.4
6/25/2008,0.4
6/26/2008,0.63
6/27/2008,0.63
6/28/2008,0.59
6/29/2008,0.57
6/30/2008,0.37
7/1/2008,0.37
7/2/2008,0.37
7/3/2008,0.37
7/4/2008,0.37
7/5/2008,0.37
7/6/2008,0.38
7/7/2008,0.38
7/8/2008,0.38
7/9/2008,0.39
7/10/2008,0.39
7/11/2008,0.39
7/12/2008,0.39
7/13/2008,0.39
7/14/2008,0.37
7/15/2008,0.37
7/16/2008,0.37
7/17/2008,0.37
7/18/2008,0.37
7/19/2008,0.37
7/20/2008,0.38
7/21/2008,0.38
7/22/2008,0.38
7/23/2008,0.39
7/24/2008,0.39
7/25/2008,0.39
7/26/2008,0.38
7/27/2008,0.37
7/28/2008,0.37
7/29/2008,0.37
7/30/2008,0.37
7/31/2008,0.37
8/1/2008,0.37
8/2/2008,0.37
8/3/2008,0.37
8/4/2008,0.37
8/5/2008,0.37
8/6/2008,0.37
8/7/2008,0.37
8/8/2008,0.37
8/9/2008,0.38
8/10/2008,0.38
8/11/2008,0.38
8/12/2008,0.38
8/13/2008,0.38
8/14/2008,0.38
8/15/2008,0.38
8/16/2008,0.38
8/17/2008,0.45
8/18/2008,0.45
8/19/2008,0.45
8/20/2008,0.45
8/21/2008,0.45
8/22/2008,0.45
8/23/2008,0.46
8/24/2008,0.46
8/25/2008,0.44
8/26/2008,0.44
8/27/2008,0.44
8/28/2008,0.44
8/29/2008,0.44
8/30/2008,0.44
8/31/2008,0.45
9/1/2008,0.45
9/2/2008,0.45
9/3/2008,0.45
9/4/2008,0.45
9/5/2008,0.45
9/6/2008,0.45
9/7/2008,0.45
9/8/2008,0.45
9/9/2008,0.45
9/10/2008,0.45
9/11/2008,0.45
9/12/2008,0.46
9/13/2008,0.46
9/14/2008,0.44
9/15/2008,0.44
9/16/2008,0.44
9/17/2008,0.44
9/18/2008,0.44
9/19/2008,0.44
9/20/2008,0.45
9/21/2008,0.45
9/22/2008,0.45
9/23/2008,0.46
9/24/2008,0.46
9/25/2008,0.45
9/26/2008,0.46
9/27/2008,0.46
9/28/2008,0.45
9/29/2008,0.45
9/30/2008,0.45
10/1/2008,0.45
10/2/2008,0.45
10/3/2008,0.45
10/4/2008,0.45
10/5/2008,0.45
10/6/2008,0.45
10/7/2008,0.45
10/8/2008,0.45
10/9/2008,0.45
10/10/2008,0.45
10/11/2008,0.45
10/12/2008,0.45
10/13/2008,0.45
10/14/2008,0.45
10/15/2008,0.45
10/16/2008,0.45
10/17/2008,0.45
10/18/2008,0.45
10/19/2008,0.45
10/20/2008,0.45
10/21/2008,0.45
10/22/2008,0.45
10/23/2008,0.45
10/24/2008,0.45
10/25/2008,0.45
10/26/2008,0.45
10/27/2008,0.45
10/28/2008,0.45
10/29/2008,0.45
10/30/2008,0.45
10/31/2008,0.45
11/1/2008,0.45
11/2/2008,0.45
11/3/2008,0.45
11/4/2008,0.45
11/5/2008,0.45
11/6/2008,0.45
11/7/2008,0.45
11/8/2008,0.45
11/9/2008,0.45
11/10/2008,0.45
11/11/2008,0.45
11/12/2008,0.45
11/13/2008,0.45
11/14/2008,0.45
11/15/2008,0.45
11/16/2008,0.48
11/17/2008,0.48
11/18/2008,0.48
11/19/2008,0.48
11/20/2008,0.48
11/21/2008,0.48
11/22/2008,0.48
11/23/2008,0.44
11/24/2008,0.44
11/25/2008,0.44
11/26/2008,0.44
11/27/2008,0.44
11/28/2008,0.44
11/29/2008,0.44
11/30/2008,0.45
12/1/2008,0.45
12/2/2008,0.45
12/3/2008,0.45
12/4/2008,0.45
12/5/2008,0.45
12/6/2008,0.45
12/7/2008,0.45
12/8/2008,0.45
12/9/2008,0.45
12/10/2008,0.45
12/11/2008,0.45
12/13/2008,0.46
12/14/2008,0.45
12/15/2008,0.45
12/16/2008,0.45
12/17/2008,0.45
12/18/2008,0.45
12/19/2008,0.45
12/20/2008,0.45
12/21/2008,0.45
12/22/2008,0.45
12/23/2008,0.45
12/24/2008,0.45
12/25/2008,0.45
12/26/2008,0.45
12/27/2008,0.45
12/28/2008,0.45
12/29/2008,0.44
12/30/2008,0.44
12/31/2008,0.44
1/1/2009,0.45
1/2/2009,0.45
1/3/2009,0.45
1/4/2009,0.45
1/5/2009,0.45
1/6/2009,0.45
1/7/2009,0.45
1/8/2009,0.45
1/9/2009,0.45
1/10/2009,0.45
1/11/2009,0.45
1/12/2009,0.45
1/13/2009,0.45
1/14/2009,0.45
1/15/2009,0.45
1/16/2009,0.45
1/17/2009,0.45
1/18/2009,0.45
1/19/2009,0.45
1/20/2009,0.45
1/21/2009,0.45
1/22/2009,0.45
1/23/2009,0.45
1/24/2009,0.45
1/25/2009,0.45
1/26/2009,0.45
1/27/2009,0.45
1/28/2009,0.45
1/29/2009,0.45
1/30/2009,0.45
1/31/2009,0.45
2/1/2009,0.45
2/2/2009,0.45
2/3/2009,0.45
2/4/2009,0.45
2/5/2009,0.46
2/6/2009,0.46
2/7/2009,0.46
2/8/2009,0.46
2/9/2009,0.45
2/10/2009,0.45
2/11/2009,0.45
2/12/2009,0.45
2/13/2009,0.45
2/14/2009,0.45
2/15/2009,0.45
2/16/2009,0.45
2/17/2009,0.45
2/18/2009,0.45
2/19/2009,0.45
2/20/2009,0.45
2/21/2009,0.45
2/22/2009,0.45
2/23/2009,0.45
2/24/2009,0.45
2/25/2009,0.45
2/26/2009,0.45
2/27/2009,0.45
2/28/2009,0.45
3/1/2009,0.45
3/2/2009,0.45
3/3/2009,0.45
3/4/2009,0.45
3/5/2009,0.45
3/6/2009,0.45
3/7/2009,0.45
3/8/2009,0.45
3/9/2009,0.45
3/10/2009,0.45
3/11/2009,0.45
3/12/2009,0.72
3/13/2009,0.72
3/14/2009,0.72
3/15/2009,0.45
3/16/2009,0.45
3/17/2009,0.45
3/18/2009,1
3/19/2009,0.84
3/20/2009,0.84
3/21/2009,0.84
How to transform that data to the profile data (used above to build control chart)? That and much more are covered by the workshop. SIGN UP!
____________________________
APENDIX: Script to install R, Shiny and Rstudio on AWS EC2 instance.
#!/bin/bash
#install R
yum install -y R
#install RStudio-Server 1.1.423-x86_64
wget https://download2.rstudio.org/rstudio-server-rhel-1.1.423-x86_64.rpm
yum install -y --nogpgcheck rstudio-server-rhel-1.1.423-x86_64.rpm
rm rstudio-server-rhel-1.1.423-x86_64.rpm
#install shiny and shiny-server (2017-08-25)
R -e "install.packages('shiny', repos='http://cran.rstudio.com/')"
wget https://download3.rstudio.org/centos5.9/x86_64/shiny-server-1.5.4.869-rh5-x86_64.rpm
yum install -y --nogpgcheck shiny-server-1.5.4.869-rh5-x86_64.rpm
rm shiny-server-1.5.4.869-rh5-x86_64.rpm
#add user(s)
useradd username
echo username:username | chpasswd
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Saturday, February 28, 2009
Real-Time Statistical Exception Detection
Does that make sense to apply statistical filtering to real-time computer performance data? I did not try as I believe analyzing last day data against historical baseline (based on dynamic statistical thresholds) would be enough to have good alert for upcoming issue and at the same time classical alerting system (based on constant thresholds, for instance, patrol or sites-scope) captures severe incidents if something completely dying.
But I see some companies do that using the following three (at least) products available on a market:
1. Integrien Alive™ (http://www.integrien.com/ )
2. Netuitive (http://netuitive.com/ )
3. ProactiveNet (now BMC), (http://documents.bmc.com/products/documents/49/13/84913/84913.pdf )
Plus Firescope http://www.firescope.com/default.htm and Managed Objects http://managedobjects.com/ do something similar)
I have recently had discussion with Integrien sales people as they did live presentation of Alive product for company I work for now.
I was impressed, it looks working good. Most interesting for me is the deference between SEDS (my approach) and their technology.
Apparently both approaches are using dynamic statistical thresholds to issue an alert.
But I think they do that using some patented complex statistical algorithms that should work well even if sample data is not normally distributed. It’s done based on some research that Dr Mazda A. Marvasti did and I am aware of this research as some of his thoughts was published in CMG (in MeasureIT) couple years ago. That consists of very good critic of SPC (Statistical Process Control) concepts applied to IT data as SPC works perfect if data is normally distributed and if not, it works not so perfect. The 1st attempt to improve SPC was MASF to regroup analyzed data and after regrouping data might be more close to normal. SEDS is based on MASF and, for instance, it looks at history in different dimension by not comparing (calculating st. deviations) hours during the same day but grouping hours by weekday and also it calculates statistic across weeks not days.
(You could find more details in my last paper. Links to some papers related to this subject including my papers can be found in this blog )
BTW In respond on his publication I did special analysis to see how far from normal the data is used by SEDS and some result of this research has been published in one of my papers. And my opinion is some data is close to normal and some still indeed is not so close and it depends of metrics, subsystems and environment (prod/non-prod) and how it’s grouped.
The key is what type of threshold the SEDS-like product uses to establish a base-line. That could be very simple – static one, or based on st. deviations, but that could be more complex thresholds such as combination of static (based on expert experiences – empiric) and simple statistical ones (based on st. deviations). SEDS uses that combination and SEDS has a several tuning parameters to tune SEDS to capture meaningful exceptions. I believe this approach is valuable (and cheap) for practical usage and several successful implementations of SEDS proves that.
But for more accurate analysis of data especially if it’s far from normal destitution, other more advanced statistical techniques could be applied and looks like this product implements that. For me it’s just another (more sophisticated) threshold calculation for base-lining. Anyway I am continue improving my approach and will be thinking about what they and others do in this area.
Other interesting observation I got from the Integrien tool live presentation:
The rate of dynamic threshold exceeding is so large that they have to put additional (static???) threshold considering that some number of exceptions are kind of normal and just a nose that should be ignored. That means if the number of exceptions is bigger that that threshold, the smart alert is issued. I did not get how this threshold is set or calculated, but it’s very high - HUNDREDS (!!!) of exceptions per interval. I believe the reason of this is they apply “Anomaly” detector to too granular data. As I stated in my last paper the better result could be reached by doing statistical after some stigmatization (SEDS does that mostly after averaging that to hourly data)
BTW SEDS uses original meta-metric to detect only meaningful exceptions (it uses EV or Exception Value - see my last paper) that allows SEDS to have fault positive rate very low.
But I see some companies do that using the following three (at least) products available on a market:
1. Integrien Alive™ (http://www.integrien.com/ )
2. Netuitive (http://netuitive.com/ )
3. ProactiveNet (now BMC), (http://documents.bmc.com/products/documents/49/13/84913/84913.pdf )
Plus Firescope http://www.firescope.com/default.htm and Managed Objects http://managedobjects.com/ do something similar)
I have recently had discussion with Integrien sales people as they did live presentation of Alive product for company I work for now.
I was impressed, it looks working good. Most interesting for me is the deference between SEDS (my approach) and their technology.
Apparently both approaches are using dynamic statistical thresholds to issue an alert.
But I think they do that using some patented complex statistical algorithms that should work well even if sample data is not normally distributed. It’s done based on some research that Dr Mazda A. Marvasti did and I am aware of this research as some of his thoughts was published in CMG (in MeasureIT) couple years ago. That consists of very good critic of SPC (Statistical Process Control) concepts applied to IT data as SPC works perfect if data is normally distributed and if not, it works not so perfect. The 1st attempt to improve SPC was MASF to regroup analyzed data and after regrouping data might be more close to normal. SEDS is based on MASF and, for instance, it looks at history in different dimension by not comparing (calculating st. deviations) hours during the same day but grouping hours by weekday and also it calculates statistic across weeks not days.
(You could find more details in my last paper. Links to some papers related to this subject including my papers can be found in this blog )
BTW In respond on his publication I did special analysis to see how far from normal the data is used by SEDS and some result of this research has been published in one of my papers. And my opinion is some data is close to normal and some still indeed is not so close and it depends of metrics, subsystems and environment (prod/non-prod) and how it’s grouped.
The key is what type of threshold the SEDS-like product uses to establish a base-line. That could be very simple – static one, or based on st. deviations, but that could be more complex thresholds such as combination of static (based on expert experiences – empiric) and simple statistical ones (based on st. deviations). SEDS uses that combination and SEDS has a several tuning parameters to tune SEDS to capture meaningful exceptions. I believe this approach is valuable (and cheap) for practical usage and several successful implementations of SEDS proves that.
But for more accurate analysis of data especially if it’s far from normal destitution, other more advanced statistical techniques could be applied and looks like this product implements that. For me it’s just another (more sophisticated) threshold calculation for base-lining. Anyway I am continue improving my approach and will be thinking about what they and others do in this area.
Other interesting observation I got from the Integrien tool live presentation:
The rate of dynamic threshold exceeding is so large that they have to put additional (static???) threshold considering that some number of exceptions are kind of normal and just a nose that should be ignored. That means if the number of exceptions is bigger that that threshold, the smart alert is issued. I did not get how this threshold is set or calculated, but it’s very high - HUNDREDS (!!!) of exceptions per interval. I believe the reason of this is they apply “Anomaly” detector to too granular data. As I stated in my last paper the better result could be reached by doing statistical after some stigmatization (SEDS does that mostly after averaging that to hourly data)
BTW SEDS uses original meta-metric to detect only meaningful exceptions (it uses EV or Exception Value - see my last paper) that allows SEDS to have fault positive rate very low.
Labels:
Integrien,
Netuitive,
ProactiveNet
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Tuesday, August 12, 2008
Exception Based Modeling and Forecasting
My new CMG paper "Exception Based Modeling and Forecasting" is accepted to be presented in the 2008 CMG conference in Las Vegas.
ABSTRACT: How often does the need arises for modeling and forecasting? Should it be done manually by ad-hoc, by project requests or automatically? What tools and techniques are best for that? When is trending forecast enough and when is a correlation with business drivers required? The answers to these questions are presented in this session. The capacity management system should automatically provide a small list of resources that needs to be modeled or forecasted; a simple spreadsheet tool can be used for that. This technique method is already implemented on the author’s environment with thousands of servers.
Here is the link to presentation slides and paper:
https://www.researchgate.net/publication/259291915_CMG2008TrubinPresent_Exception_Based_Modeling_and_Forecasting
https://www.researchgate.net/publication/221447683_Exception_Based_Modeling_and_Forecasting
The presentation is scheduled on Tuesday December 9th 2008 - Welcome!
ADDITION: This paper also scheduled to be presented in SCMG:
Raleigh, NC, October 17th 2008:
Richmond, VA, October 23d 2008:
EXAMPLE of successful modeling/forecasting from the paper:
There was the typical capacity usage problem. Some UNIX server had bad trend and there was need to make an upgrade. I was participating in the project as the Capacity Planner. I modeled several what-if scenarios based on common benchmarks (TPM) to recalculate (using just a spreadsheet formulas) UNIX box CPU usage for a few possible upgrades. The result is seen in the following Figure:
After one of my recommendations was accepted, I collected performance data on the newly upgraded configurations and the server worked as I had predicted
ABSTRACT: How often does the need arises for modeling and forecasting? Should it be done manually by ad-hoc, by project requests or automatically? What tools and techniques are best for that? When is trending forecast enough and when is a correlation with business drivers required? The answers to these questions are presented in this session. The capacity management system should automatically provide a small list of resources that needs to be modeled or forecasted; a simple spreadsheet tool can be used for that. This technique method is already implemented on the author’s environment with thousands of servers.
Here is the link to presentation slides and paper:
https://www.researchgate.net/publication/259291915_CMG2008TrubinPresent_Exception_Based_Modeling_and_Forecasting
https://www.researchgate.net/publication/221447683_Exception_Based_Modeling_and_Forecasting
The presentation is scheduled on Tuesday December 9th 2008 - Welcome!
ADDITION: This paper also scheduled to be presented in SCMG:
Raleigh, NC, October 17th 2008:
Richmond, VA, October 23d 2008:
EXAMPLE of successful modeling/forecasting from the paper:
There was the typical capacity usage problem. Some UNIX server had bad trend and there was need to make an upgrade. I was participating in the project as the Capacity Planner. I modeled several what-if scenarios based on common benchmarks (TPM) to recalculate (using just a spreadsheet formulas) UNIX box CPU usage for a few possible upgrades. The result is seen in the following Figure:
After one of my recommendations was accepted, I collected performance data on the newly upgraded configurations and the server worked as I had predicted
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Sunday, April 27, 2008
CMG'07 trip report
1. Statistical Process Control And Capacity Management (SEDS -like approach)
Igor Trubin, Ray White IBM “System Management by Exception: The Final Part”
ABSTRACT: Statistical Exception Detection System (SEDS) has been successfully used for more than seven years to automatically produce web-based exception reports and smart alerts against a performance database in a large multi-platform environment. This paper gives an overview of how SEDS uses Statistical Process Control (SPC) and Multivariate Adaptive Statistical Filtering (MASF) techniques and how it could be used as part of Lean Six Sigma. It focuses on memory usage exceptions, which SEDS captures, to proactively identify server and application performance issues.
COMMENTS:
- First time the Weekly profile (vs. daily one) Control Chart was introduced as a good source of metric report.
This paper is scheduled to be presented again in Raleigh NC SCMG meeting on May 2nd 2008: (http://regions.cmg.org/regions/scmg/spring_08/raleigh/meeting_05_02_08.htm)
Presentation: http://regions.cmg.org/regions/scmg/fall_07/richmond/SEDSCMG2007_v4.pdf2. Using SAS for Capacity Management (Vendors user group sessions)
Alla Piltser, MerilLynch – “Controlling the Bull: Managing Capacity and Performance Using SAS”
COMMENTS: That’s a Merrill Lynch Experience of providing Capacity Management for large IT shop: TeamQuest based performance monitoring and data collection infrastructure in managed UNIX, Linux, Windows and ESX environments + centralized SAS/ITRM infrastructure + exception based performance management reporting structure. There was a reference to my work as a right way to do exception based reporting.
Frank Lieble, SAS – “Bringing ITL to Life: Automating IT Capacity Management”.COMMENTS: Most interesting part of presentation is the Capacity Management Portal (ITRM based) which includes Tree-map reporting. The tool is good if there is a leak of good statisticians /sas programmers.
Peg McMahon, Justin Martin, Sprint Nextel “Death to Dashboards: Alarming, Performance Management Based on Variance, System Prioritization and Other Thoughts on Data Visualization”
ABSTRACT: When does the light on the executive dashboard turn from yellow to red? When do you order new hardware? Traditionally, these decisions are handled by setting thresholds — picking some number to use as an upper or lower limit. Thresholds might have worked well in the days of a handful of beloved systems. But for today’s complex environments, thresholding is not only painful to manage but conceptually bankrupt. Let’s talk about the problems with thresholds and dashboards and work to identify some practical alternatives. Vendors, put on your iron underwear and attend this session.
COMMENTS: The main part of this paper is just about what SEDS has been already providing and what has already been presented in my CMG papers since 2001:
“Alarming Based on Variance. The next step towards better performance monitoring is the use of a baseline approach to performance management. Using the power of statistics, the performance metrics can be analyzed to create upper and lower control limits based on the normal variance of the system’s performance. From this analysis, dynamic thresholds can be set based on the normal variance represented by the data. Implementing dynamic, variance based thresholds takes into account the system’s typical workload characteristics. Now, when a back up occurs in the middle of the night, as long as the same back up has occurred at the same time for the past several nights, the CPU threshold is not breached. In theory, an alarm will only occur when the system utilization is above or below a dynamic threshold which outlines the “normal” processing range of the system. This method of monitoring will provide a more refined approach to alarming as it will help to better identify actual performance issues. This is important when the analyst is responsible for monitoring many systems. However, when implemented across thousands of systems, there will likely be several that will have at least one hour which exceeds the variance threshold, and thus triggers alarms. When using a conventional dashboard, this improved level of monitoring creates the same problem as found earlier. How do you prioritize the order in which to resolve the performance issues? In today’s business environment the number of systems is increasing while there are fewer people to manage them. Each system has a unique impact on the business. Understanding a system’s business impact and addressing system performance issues in the correct priority will save a company significant dollars. Using a conventional stoplight dashboard for system performance management will often confuse and delay critical decision making. One way to help address the prioritization problem, using the performance variance data, could be to create a sorted list. This type of report would present the servers having the most performance variance appearing at the top. Using this list, cross-referenced with a list of systems prioritized by their business criticality would be one way to determine which problems need to be addressed first. This method is not intuitive since it requires the analyst to jog between reports. However, it is a way to use the available data in order to make the most business impacting decision. The problem in determining how to quickly prioritize system performance issues is not necessarily due to a lacking in performance data, but rather the lack of a way to properly visualize the performance data....”.
Also the paper presents another example of using a tree-map! Again, the way how tree-map can be used against performance metrics was shown in my and Lin Merritt CMG papers in 2004.
Amit Patel - “Software Performance Lifecycle at a Large National Bank”
COMMENTS: The paper shows some Statistical Process Control (SPC) technique usage. From Abstract: "… Learn how custom monitoring, Six Sigma techniques, performance testing, and daily production reports played an important role in identifying production issues…. "
To build the following control chart the “Minitab” statistical tool was used (http://www.minitab.com/)
Labels:
CMG'07 Conference
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
Monday, August 13, 2007
CMG'06: Performance Data Statistical Exceptions Analysis (Review) and my paper there...
2016 UPDATE. My paper from that year CMG conference can be found under the following link now:
SYSTEM MANAGEMENT BY EXCEPTION, PART 6

Here is the list of CMG2006 (http://www.cmg.org/) papers that discussed statistical exception detection technique:
• A Priori Evaluation of Data and Selection of Forecasting Model, Alexander Gilgur, Michael Perka MonoSphere, Inc.
LINK: http://www.daschmelzer.com/cmg2006/PDFs/038.pdf
The paper shows how important to capture “OutLaers” to produce meaningful forecasts. To do that they use some SEDS-like algorithm “Detection of Outlier Events”.
• 2006 Best Paper Award paper: Did Something Change? Using Statistical Techniques to Interpret Service and Resource Metrics. Frank M. Bereznay, Kaiser Permanente
LINK: http://cmg.org/conference/cmg2006/awards/6139.pdf
Author has some references to my papers:
“…Statistical techniques are not new to CMG. Starting in the early 1990’s there have been numerous papers addressing this subject, [Brey90], [Chu92], [Lipner92] and [Schwartz93]. This body of work seemed to cumulate with Jeff Buzen and Annie Schum’s 1995 CMG Paper introducing Multivariate Adaptive Statistical Filtering (MASF) as a new statistical technique [Buzen95]. Interest in the subject seemed to decline from that point on, with the notable exception of Igor Trubin’s set of papers on the application of MASF to many measurement and management areas [Trubin01], [Trubin02], [Trubin03], [Trubin04], and [Trubin05]. All of these papers are excellent treatments of the subject and are recommended reading…”
During discussions at this presentation some questions were asked (e.g. Sean Meidhan from BEN, who implemented some MASF ideas in there tool) about “fault positive” (fault alerts) situation sometimes generated out this technique. I had to step up and give some clarifications how SEDS handles that.
(08/2007 UPDATE:
Frank M. Bereznay have recently gave the interview to CMG MeasureIT: MeasureIT - Issue 5.08 - Getting to Know Mullen Award Winner ...
"...I also noticed that the number of papers in this area seemed to be declining since the mid to late 1990s. There were a number of papers leading up to Jeff Buzen and Annie Shum's MASF (multivariate adaptive statistical filtering) paper in 1995, and since then the trend seemed to decline with the exception of Igor Trubin's work, so I wanted to give the statistical methods some additional visibility.."
He also will be presenting at this fall SCMG meetings September 27 in Richmond and September 28 in Raleigh :
Using Statistical Techniques to Interpret Service and Resource Metrics )
• ACTIVE BASELINING IN PASSIVE DATA ENVIRONMENTS, Mike Tsykin, Fujitsu Australia, Ltd.
LINK: http://www.fujitsu.com/downloads/AU/active_baselining_in_passive_data_environments.pdf
Author in this paper uses the SPC approach for baselining. I have met with him in the previous CMG conferences discussing my SEDS technique, he picked up SPC idea and implemented that for alerting part of some Fujitsu performance tool.
• Dials for a PM Dashboard: Velocity’s Missing Twin, and Quantifying Surprise, Rich Olcott, IBM Information Technology Services
LINK: http://www.daschmelzer.com/cmg2006/PDFs/102.pdf
This paper has also some discussion how SPC should be used. E.g. how Simple Average cold be replaced by EWMA. He has references on my paper and M.Tsykin’s paper.
• My paper: SYSTEM MANAGEMENT BY EXCEPTION, PART 6, Igor Trubin, PhD
LINK: http://www.daschmelzer.com/cmg2006/PDFs/021.pdf
ABSTRACT: Statistical Exception Detection System (SEDS) has been successfully used for more than six years to automatically produce web-based exception reports against the performance data warehouse for a large, multi-platform environment. Adding some application specific metrics including middleware traffic and response times made SEDS an excellent tool for application performance management. This paper also describes how to create statistical control charts using a spreadsheet in order to capture a performance issue without using expensive tools with built-in SPC procedure.
_________________________________________________________________
CONCLUSION: There is still a big interest to MASF, SPC and SixSigma methods applied to system performance data. This year CMG'07 conference has already announced some papers related to this subject as well, including my next paper:
System Management by Exception, Part Final, Dr. Igor A. Trubin, IBM; Ray White, IBM
LINK: PDF | System Management by Exception, Part Final. - ResearchGate
ABSTRACT: Statistical Exception Detection System (SEDS) has been successfully used for more than seven years to automatically produce web-based exception reports and smart alerts against the performance data warehouse for a large, multi-platform environment. This paper starts with an overview of how SEDS uses SPC and MASF techniques and how SEDS could be used as a part of Lean/Six Sigma. Then it focuses on the memory usage exceptions that SEDS captures to proactively identify server and application performance issues.
(08/2007 UPDATE: This paper is presented also September 27 in Richmond SCMG: meting: System Management by Exception, Part Final)

SYSTEM MANAGEMENT BY EXCEPTION, PART 6
Here is the list of CMG2006 (http://www.cmg.org/) papers that discussed statistical exception detection technique:
• A Priori Evaluation of Data and Selection of Forecasting Model, Alexander Gilgur, Michael Perka MonoSphere, Inc.
LINK: http://www.daschmelzer.com/cmg2006/PDFs/038.pdf
The paper shows how important to capture “OutLaers” to produce meaningful forecasts. To do that they use some SEDS-like algorithm “Detection of Outlier Events”.
• 2006 Best Paper Award paper: Did Something Change? Using Statistical Techniques to Interpret Service and Resource Metrics. Frank M. Bereznay, Kaiser Permanente
LINK: http://cmg.org/conference/cmg2006/awards/6139.pdf
Author has some references to my papers:
“…Statistical techniques are not new to CMG. Starting in the early 1990’s there have been numerous papers addressing this subject, [Brey90], [Chu92], [Lipner92] and [Schwartz93]. This body of work seemed to cumulate with Jeff Buzen and Annie Schum’s 1995 CMG Paper introducing Multivariate Adaptive Statistical Filtering (MASF) as a new statistical technique [Buzen95]. Interest in the subject seemed to decline from that point on, with the notable exception of Igor Trubin’s set of papers on the application of MASF to many measurement and management areas [Trubin01], [Trubin02], [Trubin03], [Trubin04], and [Trubin05]. All of these papers are excellent treatments of the subject and are recommended reading…”
During discussions at this presentation some questions were asked (e.g. Sean Meidhan from BEN, who implemented some MASF ideas in there tool) about “fault positive” (fault alerts) situation sometimes generated out this technique. I had to step up and give some clarifications how SEDS handles that.
(08/2007 UPDATE:
Frank M. Bereznay have recently gave the interview to CMG MeasureIT: MeasureIT - Issue 5.08 - Getting to Know Mullen Award Winner ...
"...I also noticed that the number of papers in this area seemed to be declining since the mid to late 1990s. There were a number of papers leading up to Jeff Buzen and Annie Shum's MASF (multivariate adaptive statistical filtering) paper in 1995, and since then the trend seemed to decline with the exception of Igor Trubin's work, so I wanted to give the statistical methods some additional visibility.."
He also will be presenting at this fall SCMG meetings September 27 in Richmond and September 28 in Raleigh :
Using Statistical Techniques to Interpret Service and Resource Metrics )
• ACTIVE BASELINING IN PASSIVE DATA ENVIRONMENTS, Mike Tsykin, Fujitsu Australia, Ltd.
LINK: http://www.fujitsu.com/downloads/AU/active_baselining_in_passive_data_environments.pdf
Author in this paper uses the SPC approach for baselining. I have met with him in the previous CMG conferences discussing my SEDS technique, he picked up SPC idea and implemented that for alerting part of some Fujitsu performance tool.
• Dials for a PM Dashboard: Velocity’s Missing Twin, and Quantifying Surprise, Rich Olcott, IBM Information Technology Services
LINK: http://www.daschmelzer.com/cmg2006/PDFs/102.pdf
This paper has also some discussion how SPC should be used. E.g. how Simple Average cold be replaced by EWMA. He has references on my paper and M.Tsykin’s paper.
• My paper: SYSTEM MANAGEMENT BY EXCEPTION, PART 6, Igor Trubin, PhD
LINK: http://www.daschmelzer.com/cmg2006/PDFs/021.pdf
ABSTRACT: Statistical Exception Detection System (SEDS) has been successfully used for more than six years to automatically produce web-based exception reports against the performance data warehouse for a large, multi-platform environment. Adding some application specific metrics including middleware traffic and response times made SEDS an excellent tool for application performance management. This paper also describes how to create statistical control charts using a spreadsheet in order to capture a performance issue without using expensive tools with built-in SPC procedure.
_________________________________________________________________
CONCLUSION: There is still a big interest to MASF, SPC and SixSigma methods applied to system performance data. This year CMG'07 conference has already announced some papers related to this subject as well, including my next paper:
System Management by Exception, Part Final, Dr. Igor A. Trubin, IBM; Ray White, IBM
LINK: PDF | System Management by Exception, Part Final. - ResearchGate
ABSTRACT: Statistical Exception Detection System (SEDS) has been successfully used for more than seven years to automatically produce web-based exception reports and smart alerts against the performance data warehouse for a large, multi-platform environment. This paper starts with an overview of how SEDS uses SPC and MASF techniques and how SEDS could be used as a part of Lean/Six Sigma. Then it focuses on the memory usage exceptions that SEDS captures to proactively identify server and application performance issues.
(08/2007 UPDATE: This paper is presented also September 27 in Richmond SCMG: meting: System Management by Exception, Part Final)
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga