Popular Post

_

Showing posts with label SEDS. Show all posts
Showing posts with label SEDS. Show all posts

Tuesday, January 29, 2013

Database Space Capacity Planning: Exception Based Method

In my 2008 CMG paper "Exception based Modeling and Forecasting" I have mentioned some approach that my colleague,  coauthor and friend Ray W.  implemented to provide proactive Database Space Capacity planning:


By the way, my first and best American manager is mentioned there - Kevin M. As stated in the paper he has actually gave me the idea of the SEDS and he was the co-author of my 1st SEDS related CMG'01 paper - Exception Detection System, Based on the Statistical Process Control Concept

Tuesday, November 20, 2012

SETDS Methodology

2022 UPDATE
Some of the SETDS features are implemented into www.Perfomalist.com tool, which is described in the following post: https://www.trutechdev.com/2021/12/ and last release notes are HERE . The detailed Perfomalist CPD method is explained in this blog: https://www.trub.in/2020/08/cpd-change-points-detection-is-planed.html
_________________________________________________________________________________
Preparing my upcoming CMG'12 presentation about SEDS-lite I try to formulate what SEDS or extended version of that - SETDS actually is.

SE(T)DS is Statistical Exception (and Trend) Detection System.  It is not an application. But could be implemented by developing one. And I have done that several times (using SAS, COGNOS, BIRT, R and other programming/reporting systems). But developing SETDS-like reports/apps is just a beginning.  The most important part of SETDS is how to use that for Systems Capacity Management and how to build that in the Service Management processes. The set of my CMG papers I wrote since 2001 (list is in the very 1st post of this blog) describes that in details.

By the way it is not absolutely necessary to develop the SETDS application because starting from BMC PP and visualizer (now it is Capacity Optimizer, Perceiver and  Proactive Net) a lot of performance tools have SETDS-like features and this blog has several posts analyzing them (e.g. see Gartner's Magic Quadrant).

A Capacity Manager just need to know how to use the home made or vendor based  SETDS-like tools features efficiently and SETDS is the method. 

So bottom line is:

SETDS is the methodology of using statistical filtering, pattern recognition, active base-lining, dynamic vs. static thresholds,  IT-Control Charts, Exception Value (EV) based reporting/smart alerting and EV based  change points/trends detection to do Systems Capacity Management including Capacity Planning and Performance Engineering.

What value SETDS could bring to a company? I will formulate that later during and after my  CMG'12 presentations on which you are welcome to attend (www.CMG.org)!

(2018 UPDATE: Note , SEDS is the unsupervised SPC/MASF ML based Anomaly Detection method)

CPD perfomalist example:




Thursday, August 2, 2012

SEDS-Lite: Using Open Source Tools (R, BIRT, MySQL) to Report and Analyze Performance Data - my new CMG'12 paper

20202 UPDATE: The SEDS-Lite web app is about to be released!
_________________________________________________________
I wrote this paper with some help from Shadi G. (from Dublin, also IBMer).
The paper is based on my blog postings:
SEDS-Lite Presentation at Southern CMG Meeting in the SAS Institute
SEDS-Lite Introduction
How To Build IT-Control Chart - Use the Excel Pivot Table!
BIRT based Control Chart

HERE IS THE VIDEO PRESENTATION
Below is the abstract:
Statistical Exception Detection (SEDS) is one of the variations of learning behavior based performance analysis methodology developed, implemented and published by Author. This paper took main SEDS tools – IT-Control Chart and Exceptions (Anomalies) Detector - and showed how that could be built by Open Source type of BI tools, such as R, BIRT and MySQL or just by spreadsheet. The paper includes source codes, tool screen-shots and report input/output examples to allow reader building/developing a light version of SEDS.
-------------------------
The presentation of this paper is scheduled on December 5th, 2012 Wednesday, 2:45:00 PM - 3:45:00 PM in Las Vegas, Nevada
-------------------------

THAT IS MY SECOND CMG'12 PAPER. THE FIRST ONE ANNOUNCED HERE:

AIX frame and LPAR level Capacity Planning. User Case for Online Banking Application

Monday, May 7, 2012

SEDS-Lite Presentation at Southern CMG Meeting in the SAS Institute

Southern CMGLast Friday I have made my presentation which was announced here: SEDS-Lite: Using Open Source Tools (R, BIRT and MySQL) to Report and Analyze Performance Data. That was presented at the Southern CMG Meeting in the SAS Institute, Cary, NC. The presentation slides are linked within AGENDA and also can be downloaded from HERE


I plan to write a paper based on this presentation and to submit that to this year CMG'12 conference.


Wednesday, April 11, 2012

SEDS-Lite: Using Open Source Tools (R, BIRT and MySQL) to Report and Analyze Performance Data

My presentation with this name has been scheduled for the next Southern CMG meeting  at SAS Institute:

SCMG Meeting Raleigh
May 04, 2012

You are welcome to attend!

Thursday, April 5, 2012

Prehistory of SEDS: Virtual CMG'90 Trip Report about Control Chart Usage. Part 1.

Using the key word "Control Chart" I have found in the www.CMG.org knowledge base a few very old CMG papers with some discussions about using classical SPC approach against computer performance data.

Here is the first one:

 Fine-Grain Analysis (FGA): A Methodology for Analyzing Intermittent Performance Problems Open in a new window
  By Robert Berry & Jeffrey Hedglin 

 

The paper describes what Mainframe metrics are good to use for Control Charting. They should be two types - a. Performance Quality Measure - sounds like modern KPI... (e.g. response time);  b. System performance metrics (e.g. CPU queue length). Then the paper describes how the intermittent problem could be detected just by plotting SPC Control Charts for both type of metrics in sync (correlated).

I use that approach a lot now, but using MASF type of Control chart and specifically my IT-Control Charts.  BTW I am writing now my next CMG paper and plan to add there a couple very persuasive  examples of correlated IT-Control Charts, such as, number of concurrent user LOGONS vs. number of Ph. CPUs used by LPARS on some p770 AIX frame....

To be continued....

Friday, January 20, 2012

Control Chart usage in "Automated Analysis of Load Testing Results"


Searching again in http://academic.research.microsoft.com I have found that not only CMG papers have some discussions about anomaly detection/control charting subjects in the Systems Capacity Management field. Below are a few examples:

1. Automated Analysis of Load Testing Results , Zhen Ming Jiang published in Conference: International Symposium on Software Testing and Analysis - ISSTA , pp. 143-146, 2010


From Abstract of the paper: ".. This dissertation proposes
automated approaches to detect functional and performance
problems in a load test by mining the recorded load testing
data (execution logs and performance metrics).."

The paper has reference to three other ones (see below) related to the subject of this blog, I believe:



- I. A. Trubin and L. Merritt. Mainframe global and
workload level statistical exception detection system,
based on masf. In 2004 CMG Conference, 2004

Here is the content where my paper was referenced:
"... It is di cult for humans to interpret raw performance
metrics, as it is not clear how to categorize these raw met-
ric values into performance categories (e.g. high, medium
and low). Furthermore, some data mining algorithms (e.g.
Navie Bayes Classi er) only take discrete values as input.
We are currently exploring generic approaches to classify
performance metrics into discrete performance categories us-
ing techniques like control charts [Trubin's CMG'04 paper] to facilitate our future
work in performance analysis...."

BTW Here is a slide with MIPS control chart from that paper presentation:



2. L. Cherkasova, K. Ozonat, N. Mi, J. Symons, and
E. Smirni. Anomaly? application change? or workload
change? towards automated detection of application
performance anomaly and change. In IEEE
International Conference on Dependable Systems and
Networks, 2008.


2. B. Anton, M. Leonardo, and P. Fabrizio. Ava:
Automated interpretation of dynamically detected
anomalies. In Proceedings of the Eighteenth
International Symposium on Software Testing and
Analysis, 2009.

I plan to find and read the last two papers and maybe to report something here....



Tuesday, December 27, 2011

IT/EV-Charts as an Application Signature: CMG'11 Trip Report, Part 1


I have attended the following CMG’11 presentation (see my previous post):

A Way to Identify, Quantify and Report Change
Richard Gimarc Kiran Chennuri
CA Technologies, Inc. Aetna Life Insurance Company

Identifying change in application performance is a time consuming task. Businesses today have
hundreds of applications and each application has hundreds of metrics. How do you wade
through that mass of data to find an indication of change? This paper describes the use of an
Application Signature to identify, quantify and report change. A Signature is a compact
description of application performance that is used much like a template to judge if a change has
occurred. There are a concise set of visual indicators generated by the Signature that supports
the identification of change in a timely manner.

Here are my comments.

I like the idea of building an application characteristic called Application Signature. As described in the paper it is actually based on typical (standard) deviations of Capacity usage during the peak hours of a day.

Looking closely to the approach I see it is similar with one I have developed for SEDS but it is a bit too simplified. Anyway it is great attempt to use SEDS methodology to watch application capacity usage.

I think the weekly IT-CONTROL CHART ( see other previous post ) is a way to compare usual weekly profile with last 168 hours of data (Base-line vs. Actual), so the base-line in the format of IT-Control Charts without actual data IS AN APPLICATION SIGNATURE but in much more accurate way. It even looks like somebody’s signature:

The actual data could be significantly different, as seen below:

And that diference should be automatically captured by SEDS-like system as an exceptions and calculated how much it differs from the "Signature" using EV meta metric as a weekly sum of each hour EV values  or as a EV-Control Charts like showed here.

For instance, in this example week the application had took a bit more than 23 unusual CPU hours as calculated below:

So, if weekly EV number is 0, that means the most recently the application (server or LPAR and so on) stayed within the IT-Signature, which is GOOD – no changes happend!

The paper also shows the “calendar view“ report that consists of set of daily control charts. It is another good idea. I used to use that approach before I switched to weekly IT- charts that cover 1/4 of a month or bi-weekly ones that cover 1/2 of a month. So if you have IT-charts there is no need for the "calendar view" that sometimes is not easy to read.

Another feature could be important for capacity usage estimates: it is a balance of hourly capacity usage for the day or week vs. overall average (e.g. weekdays vs. weekends or daily “cowboy hat” profile with lunch time drop). That is supposed to be an additional IT-Signature feature. There was another CMG’11 paper that presents some interesting approach to analyze/calculate that. I plan to publish my comments about that paper. So please check my next post soon.....

Thursday, September 29, 2011

Power of Control Charts and IT-Chart Concept (Part 1)


This is the video presentation about Control Charts. It is based on my workshop I have already run a few times. It shows how to read and use Control Charts for reporting and analyzing IT systems performance (e.g. servers, applications) . My original IT-(Control) Chart concept within SEDS (Statistical Exception Detection System) is also presented.

The Part 2 will be about "How to build" control chart using R, SAS, BIRT and just 


If anybody interested I would be happy to conduct this workshop again remotely via Internet or in person. Just put a request or just a comment here.



UPDATE: See the version of this presentation with the Russian narration:

Monday, November 15, 2010

My CMG'10 presentation - "IT-Control Charts"

I will go to CMG conference this time only for one day just to present my paper "IT-Control Charts" on Wednesday December 8th 10:30 - You are WELCOME!

Check it in the CMG conference agenda  - http://www.cmg.org/cgi-bin/agenda_2010.pl?action=more&token=5030

For Russian readers (Информация по русски здесь) I made a posting about that event in my Russian mirror blog: http://ukor.blogspot.com/2010/11/cmg10_15.html

Monday, September 13, 2010

SEDS elements in the Fluke VPM (Application Performance Management tool)

I have just received 2-day training of Fluke VPM tool.   I have already mentioned in my other posting:  
"Baselining and dynamic thresholds features in Fluke and Tivoli tools" Below are my additional comments about the tool. 
  • They have the same approach as our SEDS has - to provide at the application performance status the list of business applications with most unusual response time. And they use a hit chart for that which similar SEDS used ( seeCMG'07 trip reportand the tree-map)
  • Smart alerts and the statistical filtering are used only for response time metric and the alert is issued only based on dynamic upper-limits.
  • Learning period (base-line) is "sliding" just like in main SEDS mode, but it  is based only three weeks raw data history and looks like not grouped by hour-weekdays (like SEDS does) but maybe grouped by work- and off- hours (need to check).
  • I have suggested to the VPM trainer that the statistical filtering could be applied to transaction volume metric as well and not only upper-limits, but lower-limits should be used as unusual low transaction rate needs to be be captured as a potentially bad issue.


 All in all I was impressed by the way they implemented basic SEDS principals to filter application performance metrics. I have suggested to do that in my following CMG paper in 2006  in finally that was done!  "SYSTEM MANAGEMENT BY EXCEPTION, PART 6" (can be found in the posting: 


Also that my paper suggested to use heat chart (tree-map) against network metric too: 

"...For Network devices, the bandwidth utilization can be tree-mapped. Figure shows an example of a Network tree-map. Color coding in this report could be based on exceeding constant thresholds or statistical control limits (SEDS based). Each small box represent a device (size could be indicative of relative capacity, e.g. 1 GB or 100 MB network) and a big outline box could represent a particular application or site (e.g. building)..."



Wednesday, October 21, 2009

Lower Control Limit Usage Examples for IT Capaciy Management

I have recently posted the following question as LinkedIn discussion subject for "Statistical Process Control" group: "Does it make any sense to use Control Charts for capacity management?" and got one pessimistic comment, which included the following statement:

"...The only situation I can think of using a control chart for capacity is if you had a piece of equipment that if over utilized would cause damage or premature wear in which case you would only have an upper control..."

I disagree. My system (SEDS) has a special part (updated lists) called "Unusual Capacity Usage OUTSIDERS" that can help to capture some serious issues with servers, such as database going down, LPAR migration out of a host and other unusual capacity releases, that  are not necessarily good things:

The following control charts from my up-coming CMG'09 workshop presentation are good illustrations of those type of finding SEDS captures:

1. Vmware host issue (VM migration):



2. Unisys server database is down:




3. Mainframe application unusual low CPU usage:


Sunday, September 20, 2009

Near-Real-Time IT-Control Charts

On the next Thursday September 24, 2009 in the Richmond's SCMG meeting I am going to present my updated version of previous presentation called "Power of Control Chart". This time the focus is on Near-Real-Time IT-Control Charts. Below is the clip that shows the example of Near-Real-Time IT-Control Chart simulated by R-program:

The presentation will be published in SCMG site: http://regions.cmg.org/regions/scmg/fall_09/richmond/meeting_09_24_09b.htm

Thursday, July 23, 2009

Real-Time Control Charts for SEDS

I still analyze different tools that capture computer application abnormalities based on real-time data. In addition to Integrien (now it is a part of VMware's tool called AliveVMand Netuitive I have recently looked at BMC ProActive Net Analytics. I have spoken with BMC SMEs and they showed me a live demo of the tool. I always respected BMC (and espetialy BGS) as actually the inventor of this approach (MASF) and long ago I used to analyze statistical exceptions using BMC Visualizer and BMC Perceive (BTW I have published in my papers a few examples how I did that) . Now they have another and very good tool for the same purpose (http://documents.bmc.com/products/documents/49/13/84913/84913.pdf)

Watching the live presentation I got a positive impression of how that works for complex applications and transactions correlating different abnormal events with possibility to reduce false positives situations. Interesting that the combination of dynamic and static thresholds are used there to generate alarms. Just like SEDS does - static one to capture hot issues (run-aways and leaks) and statistical ones for early warnings.

Now I have a very difficult task to choose from those three products (plus SEDS) to recommend to my management...

Speaking about SEDS, I have decided to play with near-real time data to see how difficult would be to redesign SEDS making it works more similar with mentioned above modern and serious tools. Fortunately SEDS is just a bunch of SAS Marcos with parameters which helped me to make the adjustment needed to include today's data. And surprisingly that was pretty easy task! I spent only a couple days to developed a "real-time SEDS" prototype. Currently what it only does is building every hour the real-time Control Charts that can be seen at the beginning of this post.

I plan to include some details about real-time Control Charts to my upcoming CMG'09 Workshop.

Thursday, June 7, 2007

System Management by Exception

Greetings!

To keep the discussion about how to Manage computer Systems by Exception (e.g. by  using SPC, APC, MASF, 6-SIGMA, SETDS and other techniques), I run this blog and also publish/present white papers at the www.CMG.org.  Please take a look at the following set of CMG papers related to Statistical Exception Detection System (SEDS or SETDS):

2017 -  The Model Factory - Correlating Server and Database Utilization with Customer Activity"

2016 - Is your Capacity available? 

2012 - SEDS-Lite:  Using Open Source Tools (R, BIRT and MySQL) to Report and Analyze Performance Data 

2008 Exception Based Modeling and Forecasting

2005 - Capturing Workload Pathology by Statistical Exception Detection System 

2004 - Mainframe Global and Workload Level Statistical Exception Detection System Based on MASF

2003 - Disk Subsystem Capacity Management Based on Business Drivers I/O Performance Metrics and MASF

2002 - Global and Application Levels Exception Detection System, Based on MASF Technique

2001 - Exception Detection System, Based on the Statistical Process Control Concept