Popular Post

_

Showing posts with label Capacity Managment. Show all posts
Showing posts with label Capacity Managment. Show all posts

Tuesday, December 27, 2011

IT/EV-Charts as an Application Signature: CMG'11 Trip Report, Part 1


I have attended the following CMG’11 presentation (see my previous post):

A Way to Identify, Quantify and Report Change
Richard Gimarc Kiran Chennuri
CA Technologies, Inc. Aetna Life Insurance Company

Identifying change in application performance is a time consuming task. Businesses today have
hundreds of applications and each application has hundreds of metrics. How do you wade
through that mass of data to find an indication of change? This paper describes the use of an
Application Signature to identify, quantify and report change. A Signature is a compact
description of application performance that is used much like a template to judge if a change has
occurred. There are a concise set of visual indicators generated by the Signature that supports
the identification of change in a timely manner.

Here are my comments.

I like the idea of building an application characteristic called Application Signature. As described in the paper it is actually based on typical (standard) deviations of Capacity usage during the peak hours of a day.

Looking closely to the approach I see it is similar with one I have developed for SEDS but it is a bit too simplified. Anyway it is great attempt to use SEDS methodology to watch application capacity usage.

I think the weekly IT-CONTROL CHART ( see other previous post ) is a way to compare usual weekly profile with last 168 hours of data (Base-line vs. Actual), so the base-line in the format of IT-Control Charts without actual data IS AN APPLICATION SIGNATURE but in much more accurate way. It even looks like somebody’s signature:

The actual data could be significantly different, as seen below:

And that diference should be automatically captured by SEDS-like system as an exceptions and calculated how much it differs from the "Signature" using EV meta metric as a weekly sum of each hour EV values  or as a EV-Control Charts like showed here.

For instance, in this example week the application had took a bit more than 23 unusual CPU hours as calculated below:

So, if weekly EV number is 0, that means the most recently the application (server or LPAR and so on) stayed within the IT-Signature, which is GOOD – no changes happend!

The paper also shows the “calendar view“ report that consists of set of daily control charts. It is another good idea. I used to use that approach before I switched to weekly IT- charts that cover 1/4 of a month or bi-weekly ones that cover 1/2 of a month. So if you have IT-charts there is no need for the "calendar view" that sometimes is not easy to read.

Another feature could be important for capacity usage estimates: it is a balance of hourly capacity usage for the day or week vs. overall average (e.g. weekdays vs. weekends or daily “cowboy hat” profile with lunch time drop). That is supposed to be an additional IT-Signature feature. There was another CMG’11 paper that presents some interesting approach to analyze/calculate that. I plan to publish my comments about that paper. So please check my next post soon.....

Thursday, February 17, 2011

Jonathan Gladstone: Threshold Management Diagram

Jonathan Gladstone has worked with a team to implement pro-active Mainframe CPU usage monitoring, basing his design partly on presentations and conversations with Igor Trubin (currently of IBM) and Boris Ginis (of BMC Software).

His system does not generate any alerts on this basis, but it’s a good place to go to
  • find out what’s been running hot (or cool) at the system level, and/or
  • figure out why at the service class level.

It compares each interval (in this case every 10 minutes) of the most recent day’s utilization (by system and by service class) with the average for a given hour on a given day of the week over the past six weeks. Each interval is compared to the set of the last 36 values in a similar timeframe. If more than one interval in an hour is higher than the 98th percentile for its hour & day, the hour is marked yellow; if more than four intervals are high the hour is marked red. If more than one interval is lower than the 2nd percentile for its hour & day, the hour is marked blue. Anything in between (i.e. anything that falls within roughly x-bar±2SD) is green.

Here’s the main “CPU Overview” page from his system:



The thumbnails give an idea of what’s going on – green is within normal range. Let’s look at the Sunday, Jan. 23rd (just because all the colours are there). Clicking on any thumbnail shows that day close up:



Without going into details about what runs in which systems, we can see that they’re listed in reverse alpha order and, of course, anyone who’s looking at this knows which system is which. The user can see that a lot of systems were running well below their normal utilization on this particular Sunday. That’s mostly because of some special testing: our developers were asked to stay off the systems if they could. To see more detail let’s choose SCA6, which has all of the colours. If we click anywhere on the bar for SCA6, this next level of detail is shown:


  

That chart shows the system’s total utilization (from SMF70s) for individual 10-minute intervals (green area) compared to the average, high (98%ile) and low (2%ile) values for each hour based on the last six weeks. We see why some hours are marked red, yellow or blue instead of green according to the rules above. Clicking  anywhere on the green area gets a long page full of control charts that show the same information for each defined service class within that system (from SMF72s).

Among them the following, BATH_A6, is high-priority batch. Clearly it was driving some of the yellow and red flags for this system in the 2-3 and 5-7h windows:





(This post is published here with Jonathan’s Gladstone permission. He retains all publication rights and copyright for this material)

Monday, December 13, 2010

Video report about my 1-day attending/presenting at CMG'10 Conference in Orlando


MyCMG'10 presentation is described here:
http://itrubin.blogspot.com/2010/11/my-cmg10-presentation-it-control-charts.html

Here is a picture me siting in anther CMG'10 session:
https://www.facebook.com/photo.php?fbid=10150196216458678&set=a.10150196216138678.334041.120810323677&type=1&ref=nf

Thursday, May 20, 2010

Baselining and dynamic thresholds features in Fluke and Tivoli tools

I have just attended two demo sessions about Fluke VPM and Tivoli Monitoring tools and both presentations included the following new features related to performance Exception Detection technology:

1. Fluke (NetFlow Tracker - http://www.skomplekt.com/pdf/PerformanceAndScalability(fnet).pdf)
- Baselined alarms trigger when normal usage is exceeded.
- Automatically choose a threshold using a baseline, or manually specify.
- Tracker baselines individual elements of a report, not just the total.
- Baseline can be static (learn once) or update weekly (learn every week).

2. Tivoli Monitoring Version 6.2.2 ( http://publib.boulder.ibm.com/infocenter/tivihelp/v15r1/index.jsp?topic=/com.ibm.itm.doc_6.2.2/new_version622.htm)
- The bar chart, plot chart, and area chart have a new "Add Monitored Baseline" tool for selecting a situation to compare with the current sampling and anticipated values. The plot chart and area chart also have a new "Add Statistical Baseline" tool with statistical functions. In addition, the plot chart has a new Add Historical Baseline tool for comparing current samplings with a historical period.
- Situation overrides for dynamic thresholding.

That is interesting as I personally discussed with Tivoli team possibility to add SEDS like features to Tivoli tools just before I left IBM (about 3 years ago). It looks like they finally have implemented some elements of that technology!

I will be playing with all those features soon and plan to add more comment about how that works.

Tuesday, December 29, 2009

Exception Value (EV) and OPNET Panorama

I have recently looked at the following OPNET resources to get impressions of OPNET Panorama tool:

1. Link to website: http://www.opnet.com/solutions/application_performance/panorama.html
2. White paper downloaded from that site: "Understanding OPNET Panorama’s Performance Analysis Engines"

General comment: OPNET becomes the next generation tool provider that has “learning behavior” capabilities that are similar with what I do for years with my SEDS and with tools from other Vendors, such as Netuitive, Integrien and ProactiveNet (BMC) that I have recently studied (check my older postings: http://itrubin.blogspot.com/2009/02/realtime-statistical-exception.html).

Special comment: In the OPNET white paper I read: "Metrics that exhibit deviations from normal are automatically identified and assigned scores based on “how abnormal” their behavior is."
This is very close to what I introduced in my 1st CMG paper in 2001("Exception Detection System, Based on the Statistical Process Control Concept" ) and called ExtraVolume of a metric (I call that now Exception Value (EV) meta-metric). In OPNET’s white paper referenced to that my CMG’01 paper, but they did not mention that they use very similar approach to rang exceptions (“Area Out vs. Limit Range In Metric Scoring”).
Here is example from my 1st CMG paper of the usage EV metric to build the TOP exceptional Unix servers list:


I had even tied to normalize that metric to some Unix benchmarks (TPC) to compare ranges of exceptional capacity usage ecross diferent type of servers and configurations. The example of the report is in the paper.

Al in all, it is a good news that another vendor has been adopting that technology (maybe with some of my work’s influence!). Based on my experience with OPNET tools (very limited, just ITD Guru for network and some server behavior simulations), that tool most likely can be trusted. To speak more I need at least to play with demo….

Wednesday, October 21, 2009

Lower Control Limit Usage Examples for IT Capaciy Management

I have recently posted the following question as LinkedIn discussion subject for "Statistical Process Control" group: "Does it make any sense to use Control Charts for capacity management?" and got one pessimistic comment, which included the following statement:

"...The only situation I can think of using a control chart for capacity is if you had a piece of equipment that if over utilized would cause damage or premature wear in which case you would only have an upper control..."

I disagree. My system (SEDS) has a special part (updated lists) called "Unusual Capacity Usage OUTSIDERS" that can help to capture some serious issues with servers, such as database going down, LPAR migration out of a host and other unusual capacity releases, that  are not necessarily good things:

The following control charts from my up-coming CMG'09 workshop presentation are good illustrations of those type of finding SEDS captures:

1. Vmware host issue (VM migration):



2. Unisys server database is down:




3. Mainframe application unusual low CPU usage:


Sunday, September 20, 2009

Near-Real-Time IT-Control Charts

On the next Thursday September 24, 2009 in the Richmond's SCMG meeting I am going to present my updated version of previous presentation called "Power of Control Chart". This time the focus is on Near-Real-Time IT-Control Charts. Below is the clip that shows the example of Near-Real-Time IT-Control Chart simulated by R-program:

The presentation will be published in SCMG site: http://regions.cmg.org/regions/scmg/fall_09/richmond/meeting_09_24_09b.htm

Thursday, June 7, 2007

System Management by Exception

Greetings!

To keep the discussion about how to Manage computer Systems by Exception (e.g. by  using SPC, APC, MASF, 6-SIGMA, SETDS and other techniques), I run this blog and also publish/present white papers at the www.CMG.org.  Please take a look at the following set of CMG papers related to Statistical Exception Detection System (SEDS or SETDS):

2017 -  The Model Factory - Correlating Server and Database Utilization with Customer Activity"

2016 - Is your Capacity available? 

2012 - SEDS-Lite:  Using Open Source Tools (R, BIRT and MySQL) to Report and Analyze Performance Data 

2008 Exception Based Modeling and Forecasting

2005 - Capturing Workload Pathology by Statistical Exception Detection System 

2004 - Mainframe Global and Workload Level Statistical Exception Detection System Based on MASF

2003 - Disk Subsystem Capacity Management Based on Business Drivers I/O Performance Metrics and MASF

2002 - Global and Application Levels Exception Detection System, Based on MASF Technique

2001 - Exception Detection System, Based on the Statistical Process Control Concept