Popular Post

_

Wednesday, February 15, 2012

Gartner's Magic Quadrant for Application Performance Monitoring and Behavior Learning Engine


I strongly believe that my SEDS or SETDS (Statistical Exception and Trend Detection System) could be treated as BLE – Behavior Learning Engine. SEDS or SETDS (new name I use now) is not recognized by the following Gartner’s research, but BLE is. 

Gartner 2011 research (G00215740) called “Magic Quadrant for Application Performance Monitoring” ( can be downloaded here) admitted that one of the important functionality dimensions of APM is “Applications Performance Analytics” which descried in the research and can be seen in below quotes:
 
That includes BLE which is indeed the essential component of Application Performance Analytics. The research includes the several Vendors/tools analyses that showed in the Quadrant picture below:

 

But only the following vendors were indicated in the research as having tools with strong behavior learning features: 

ASG
BMC software
 


CA Technologies
HP
Compuware
 
 IBM
 
 

I hope the SETDS implementation offering withing the IBM consulting service (which I currently do) could shift the company in that Magic Quadrant to the right...

Saturday, February 4, 2012

FiOS Upload Speed Measuring - Do not Trust Verizon Speed test!


Still having upload problem with my FiOS Internet, I have been advised by my friend to test and graph the upload process using the following website:  http://www.connectionanalyzer.com/home.php - 

I did that and here are some results that shows the problem clearly - very long pauses/delays up to 1-2 sec in the middle of transfers! That averages the speed to very low number. Between pauses the speed is good.

test 1:

test 2:

test 3:
Last one is the worst and the overall averaged measurement was close to my observation I published in the previous posting ~ 0.18 Mbps upload speed!:

After the test I used Verizon speed test which still shows the speed I actually pay for (25/25):

  
Amazing! That why I had a problem to make Verizon to pay attention to my problem! Based on their test I do not have any problem!!!!

Anyway I convinced Verizon (maybe because I publishing issue update on my blogs...) to send me a technician to fix the problem at my house. I will publish other test after the issue is gone.


Tuesday, January 31, 2012

How to get the actual spreadsheets I have developed to build control charts.

I got that question from one of my blog's reader. Here is my answer to him:

I share my spreadsheets only during my live presentations I occasionally do (CMG.org local and international ones); so you can keep checking (on my blog) when my next one will be and you are welcome. I plan to submit my next CMG paper for this year International Conference in Las Vegas, so try to attend and we will talk.  You can invite me to you local CMG chapter! If you do not have one - create that! Ask me how to do that and I can help.

If you cannot reach me that way, I usually do very detailed explanation how my spreadsheets (like control chart builder) work on my blog and papers, so I believe it is possible just to recreate them. Let me know if you have any specific question what is not clear in my postings or papers and I will try to give you more explanations.

And finally I am IBM consultant! Your company can call me to help you implement my ideas!

Sunday, January 29, 2012

Bad and Good FiOS Upload Comparison

This is the continuation of previous research published here:
 FiOS Problem: Large File Upload Speed Analysis

My son repeated the same test I did but on his own laptop and got the same and even more interesting result that shows the problem very clear: it does the fast upload at 2.5 Mgps during transferring 1st 2-3 Mg and then it degrades to ~0.17 Mgps for the rest of file traansfer with some ocational spiks: 



Then my son went to his next door friend who has very fast 35/25 FiOS  and tested the upload speed there the same way we did from my house. Everything was extremely fast as seen below:
Them I went to my other neighbor (same 35/25 FiOS) with my laptop, ran my test  and got also the very good result:
That was a 200Mb file upload to YouTube and it shows actual upload speed about 25% out of 100Mbps which  matches the standard speed test. I even tested sending attached 2 Mg file via my Verizon.net account - NO PROBLEM:

So we made the clear experiment and now I am sure there is no problem with my laptop and the FiOS in our neighborhood is OK. The problem is around my house. I am sending this report to Verizon support. We will see how they help me. The update will be in the next additional posts in my personal blog: http://trubinigor.blogspot.com.



Friday, January 27, 2012

FiOS Problem: Large File Upload Speed Analysis

I am afraid my home internet (Verizon FiOS) does not provide the upload speed that actually I pay for. I have 15/5 Mbps, so upload speed should be about 5 Mbps.

Speed test result (http://www.speakeasy.net/speedtest/) gives the following:
Good, right? But that test uses very small size file, I believe. So I could not resist to make some experiment to measure the real upload speed I have for relatively large file (~8Mb). 
See result here: 
 
That means I have about 0.17% out of 100Mbps = 0.17 Mbps instead of 5Mbps!!!!

That is interesting... a small file upload is fast, so standard speed test is not capturing a problem, but long upload is degrading significantly! 

It cannot be a trick from Verizon to hide their problems I hope, that should be some network defect or capacity issue.

Sometimes I see errors:
- standard speed test from some distant locations shows error (see example below from Seattle): 
 - Plus when I am attaching >2 Mb file in my e-mail it returns error after uploading less than half of file :


Note: Ironically Verizon "home agent" program senses this e-mail problem, but cannot help at all!

So it is the  real problem for a blogger!

I have opened a ticket with Verizon and having a battle right now with them to fix that. See details: http://trubinigor.blogspot.com/2012/01/facebook-verizon-account-responded-on.html 


Monday, January 23, 2012

Quantifying Imbalance in Computer Systems: CMG'11 Trip Report, Part 2

UPDATE 2018:
The technique was successfully tested in the SonR (SEDS based Anomaly detection system) as described in the following post:

"My talk, "Catching Anomaly and Normality in Cloud by Neural Net and Entropy Calculation", has been selected for #CMGimpact 2019

_______________________________________________________  original post:
As I promised in CMG'11 Trip Report, Part 1 here is my comments and some follow up analysis of the following paper: Quantifying Imbalance in Computer Systems that was written and presented at CMG'11 by Charles Loboz from Windows Azure.

The  idea is to calculate imbalance of a system by using an entropy property which well know in the physics , economics and in the information theory

In my other past posting I rose the following question:
 "can the information theory (entropy analysis) could be applied to performance exception detection?"

Looks like the idea from  the mentioning CMG paper of using entropy calculation against system performance data could lead to the answer of that my question!

 Here is the quote from the paper: 



"...Theil index is based on entropy - it describes the excess entropy in a system. For a data set xi,
i=1..n the Theil index is given by:

where n is the number of elements in the data set and xavg is the average value of all elements in the data set. To underline the application of the Theil index to measure  imbalance in computer systems we call it henceforth the Imbalance Coefficient (IC). 

Examining closer the IC formula above we can derive several properties:
  • (1) the ratio xi/xavg describes how much element i is above or below the average for the whole set. Thus IC involves only the ratio of each element against the average, not the absolute values of theelements.
  • (2) IC is dimensionless .– thus allows to compare imbalance between sets of substantially different quantities, for example when one set contains disk utilizations and another disk response times.
  • (3) The minimum value of IC is zero - when all elements of the data set are identical. The maximum value of the Imbalance Coefficient is log(n) - when all elements but one are equal; the maximum IC depends thus on the set size.
  • (4) We can view Imbalance Coefficient as a description of how concentrated is the use of some resource .– large values mean fewer users use most of the resource, small values mean more equal sharing.

We also define, for convenience, Normalized Imbalance Coefficient (nIC) as

to account for both imbalance within the set and the maximum entropy in that set. The nIC value ranges from 0 to 1 thus enabling comparison of imbalance between data sets with differing number of elements..."

Author applied that to the multiple disks utilization analysis, but he mentioned that approach could be used for measuring other computer subsystems imbalance. So I decided to try to calculate the imbalance of CPU utilization during the day (24 hours) and a week (168 hours) because the  imbalance of capacity usage during a day or week is a pretty common concern. Also using my way to group base-line vs. actual data I have applied that twice to compare an "average" weekly/daily utilization vs. last week/days of actual utilization.

The raw data is the same as for the last Control Charting exercise I published here in the  series of posts ( see EV-Control Chart as an example), where the actual data (in black) vs. historical averages (in green) are shown below:

Here is the result of calculating the actual vs. averaged nIC Imbalance difference for all 168 hours and for each weekdays (7 days by 24 hours):


You can see that in the day when the anomaly of CPU usage started - Wednesday - the imbalance was significantly different and all in all weekly imbalance was significantly different too!  So indeed that metric can be use to capture some performance metric anomalies (pattern changes). 

FYI: Here is the spreadsheet snapshot with actual calculation I used: 

How better that method of imbalance change checking to compare with more traditional ways to do that (e.g. based on deviations) is hard to say. My personal preference is still EV-concept. Anyway someone needs to try that against more data...

BTW I have found another paper which relates to that topic:

Quantifying Load Imbalance on Virtualized Enterprise Servers by
Emmanuel Arzuaga and David R. Kaeli

In that paper here is the clear statement about imbalance: "A typical imbalance metric based on the resource utilization of physical servers is the standard deviation of the CPU utilization".

Still an entropy is interesting system property that should give us additional good source of information for pattern recognition, I believe. For instance, the balance of Capacity usage of large frames with a lot of LPARS (AIX p7s or VMware hosts)  could be monitored by using that nIC metric to apply some possibly an automatic way to rebalanced capacity usage by using partition mobility or v-motion technologies.   

Friday, January 20, 2012

LinkedIn Discussion: "How to write a book or blog”

I have responded on the LinkedIn Discussion: " How to write a book or blog  initiated by professional blogger Greg SchulzAnd I have got the following  excellent advises I am going to follow: 

"Igor with all of your white papers and posts, you probably have a good basis for a book or ebook. Likewise, in the course of doing a book project, there tends to be a lot of content that ends up on the "cutting room floor" that makes for future blogs posts, articles, tips, etc.

Sounds like a good theme topic for a book, particular if you took an angle of "...past, present and future...". The idea of the past, present and future is to discuss how statistical and empirical measurements have evolved, are being used and will continue to be important in the future. After all, you (or your cloud provider) cannot effectively manage what they do not have insight or awareness into. Hence the importance and role of statistical and empirical analysis.

Of course, you can play the buzzword bingo game angle by working in how big data and hadoop tie into the supporting statistical analysis. Try an experiment assuming that you have stats enabled for your websites, which is look at normal traffic patterns. Then do a post with a title along the lines of "statistical monitoring with big data" and see what changes in traffic patterns occur.
...
I have one primary blog (e.g. http://storageioblog.com ) where either most of my material goes initially or as a follow-up to items that appear elsewhere. Now that I think about it, I guess I do have other blogs that either pick up my feeds automatically, or that I periodic visit and quickly cross post if wordpress friendly. There are also a bunch of other sites where articles, topics, pod casts, videos or guest posts appear in addition to those that syndicate my blog feed (e.g. via http://storageioblog.com/RSSfull.xml orhttp://storageioblog.com/RSSfullArchive.xml ).

My RSS feeds are free to anyone to use as long as they retain links that are in the post maintain attributions and copyrights do not insert content or posts from others in-line of a post, or otherwise change the content context. Likewise, sites are free to use excerpts as long as they attribute back to the source and preserve copyrights including if/when put into creative commons...."


Control Chart usage in "Automated Analysis of Load Testing Results"


Searching again in http://academic.research.microsoft.com I have found that not only CMG papers have some discussions about anomaly detection/control charting subjects in the Systems Capacity Management field. Below are a few examples:

1. Automated Analysis of Load Testing Results , Zhen Ming Jiang published in Conference: International Symposium on Software Testing and Analysis - ISSTA , pp. 143-146, 2010


From Abstract of the paper: ".. This dissertation proposes
automated approaches to detect functional and performance
problems in a load test by mining the recorded load testing
data (execution logs and performance metrics).."

The paper has reference to three other ones (see below) related to the subject of this blog, I believe:



- I. A. Trubin and L. Merritt. Mainframe global and
workload level statistical exception detection system,
based on masf. In 2004 CMG Conference, 2004

Here is the content where my paper was referenced:
"... It is di cult for humans to interpret raw performance
metrics, as it is not clear how to categorize these raw met-
ric values into performance categories (e.g. high, medium
and low). Furthermore, some data mining algorithms (e.g.
Navie Bayes Classi er) only take discrete values as input.
We are currently exploring generic approaches to classify
performance metrics into discrete performance categories us-
ing techniques like control charts [Trubin's CMG'04 paper] to facilitate our future
work in performance analysis...."

BTW Here is a slide with MIPS control chart from that paper presentation:



2. L. Cherkasova, K. Ozonat, N. Mi, J. Symons, and
E. Smirni. Anomaly? application change? or workload
change? towards automated detection of application
performance anomaly and change. In IEEE
International Conference on Dependable Systems and
Networks, 2008.


2. B. Anton, M. Leonardo, and P. Fabrizio. Ava:
Automated interpretation of dynamically detected
anomalies. In Proceedings of the Eighteenth
International Symposium on Software Testing and
Analysis, 2009.

I plan to find and read the last two papers and maybe to report something here....



Tuesday, December 27, 2011

IT/EV-Charts as an Application Signature: CMG'11 Trip Report, Part 1


I have attended the following CMG’11 presentation (see my previous post):

A Way to Identify, Quantify and Report Change
Richard Gimarc Kiran Chennuri
CA Technologies, Inc. Aetna Life Insurance Company

Identifying change in application performance is a time consuming task. Businesses today have
hundreds of applications and each application has hundreds of metrics. How do you wade
through that mass of data to find an indication of change? This paper describes the use of an
Application Signature to identify, quantify and report change. A Signature is a compact
description of application performance that is used much like a template to judge if a change has
occurred. There are a concise set of visual indicators generated by the Signature that supports
the identification of change in a timely manner.

Here are my comments.

I like the idea of building an application characteristic called Application Signature. As described in the paper it is actually based on typical (standard) deviations of Capacity usage during the peak hours of a day.

Looking closely to the approach I see it is similar with one I have developed for SEDS but it is a bit too simplified. Anyway it is great attempt to use SEDS methodology to watch application capacity usage.

I think the weekly IT-CONTROL CHART ( see other previous post ) is a way to compare usual weekly profile with last 168 hours of data (Base-line vs. Actual), so the base-line in the format of IT-Control Charts without actual data IS AN APPLICATION SIGNATURE but in much more accurate way. It even looks like somebody’s signature:

The actual data could be significantly different, as seen below:

And that diference should be automatically captured by SEDS-like system as an exceptions and calculated how much it differs from the "Signature" using EV meta metric as a weekly sum of each hour EV values  or as a EV-Control Charts like showed here.

For instance, in this example week the application had took a bit more than 23 unusual CPU hours as calculated below:

So, if weekly EV number is 0, that means the most recently the application (server or LPAR and so on) stayed within the IT-Signature, which is GOOD – no changes happend!

The paper also shows the “calendar view“ report that consists of set of daily control charts. It is another good idea. I used to use that approach before I switched to weekly IT- charts that cover 1/4 of a month or bi-weekly ones that cover 1/2 of a month. So if you have IT-charts there is no need for the "calendar view" that sometimes is not easy to read.

Another feature could be important for capacity usage estimates: it is a balance of hourly capacity usage for the day or week vs. overall average (e.g. weekdays vs. weekends or daily “cowboy hat” profile with lunch time drop). That is supposed to be an additional IT-Signature feature. There was another CMG’11 paper that presents some interesting approach to analyze/calculate that. I plan to publish my comments about that paper. So please check my next post soon.....

Tuesday, December 6, 2011

Application Signature: some of my SEDS ideas are at work

I am at CMG'11 conference now (in DC) presenting nothing this year (1st time for the last 11 years!), but I enjoy the conference and especially when my work is referenced.

Here is the example from paper called "Application Signature: A Way to Identify, Quantify and Report Change" which s presenting today at 4 pm by Richard Gimarc from CA Technologies, Inc and Kiran Chennuri from Aetna Life Insurance Company:

'...We readily admit that we are “standing on the shoulders of giants”; leveraging the work of others in the field to develop our own interpretation, implementation and use of an Application Signature....
... Perhaps the most influential work is by Igor Trubin. Starting in 2001, Trubin built on the ideas proposed by Buzen and Shum to develop the Statistical Exception Detection System (SEDS). Basically, SEDS “is used for automatically scanning through large volumes of performance data and identifying measurements of global metrics that differ significantly from their expected values”. Again, we see common ground with our use of an Application Signature. The points we leverage from Trubin’s work are:
  • Identify when performance metrics exceed of fall below expectation
  • Note and record the exceptions
  • Estimate the size of each exception rather than just recording its occurrence
  • Use control charts as a visual tool for examining current performance versus expected performance
 ...
What do you do when a change is identified?
  • Quantify the change. Does your current measurement exceed the Signature by 5%, or 100%? We are considering implementing a technique similar to what was described by Trubin.
  • Grade the change as either good or bad. If a metric increases, is that an indication of a bad change? Not always. Consider workload throughput; an increase in workload throughput is probably a good change. We need to find a way to customize each Application Signature metric to recognize and highlight both good and bad changes.
  • Develop a historical record of changes. Again, this is an idea developed by Trubin. A historical record will provide the application development and support staff with a quantitative description of sensitive application characteristics that may warrant improvement. 
...'
Some other anthers' work are referenced. I need to read that carefully and will report here about that in the other posts. Looking forward to attend that presentation! 

Richard and Kiran, thank you for referencing my work!



Tuesday, November 29, 2011

Finding the Edge of Surprise by Rich Olcott

I have definitely overlooked the following very good article of my CMG and IBM acquaintance:

MeasureIT - Issue 5.03 - Finding the Edge of Surprise by Rich Olcott 

At the 1st glance that article has a good overview of Classical SPC with some original suggestion how to apply that to IT data. Also I like the name of the article which could be a good short and metaphoric description of the main topic of this entire blog! 

BTW He provided there the reference to my CMG'2004 paper: “Mainframe Global and Workload Levels – Statistical Exception Detection System, Based on MASF,” CMG Proceedings (2004). The link to that my paper is published on very 1st posting of this blog!

And I have already mentioned  his previous work at my other posting: 
Aug 13, 2007
Dials for a PM Dashboard: Velocity's Missing Twin, and Quantifying Surprise, Rich Olcott
I plan to reread both his works and to add more comments-thoughts....

Wednesday, November 9, 2011

SEDS-Lite: Using Open Source Tools (R, BIRT and MySQL) to Report and Analyze Performance Data

Last Thursday we had a very good Southern Computer Measurement Group meeting of 16 attendees in Richmond VA, where I have presented the material about how to use R, BIRT, MySQL and EXCEL to analyze and report systems' performance data having as an example some real Unix server CPU utilization data for control charting.

Agenda is still on SCMG website and now my presentation slides are published and linked there:

SEDS-Lite: Using Open Source Tools (R, BIRT and MySQL) to Report and Analyze Performance Data
(slides).



Tuesday, October 11, 2011

My Southern CMG Presentation in Richmond Is About Open Source Tools for Capacity Management

I have been invited to make my new presentation on the 2011 Fall SCMG Meeting. See agenda here.


My presentation will be actually a compilations of some of my last posts in this blog: 

UCL=LCL : How many standard deviations do we use for Control Charting? Use ZERO! 
BIRT based Control Chart 
One Example of BIRT Data Cubes Usage for Performance Data Analysis 
How To Build IT-Control Chart - Use the Excel Pivot Table! 
Power of Control Charts and IT-Chart Concept (Part 1) 
Building IT-Control Chart by BIRT against Data from the MySQL Database 
- EV-Control Chart


So please plan to attend ! (registration is here)

Monday, October 10, 2011

Is Anomaly Detection Similar to Exception Detection? Apply SEDS for Information Security!

Sometimes I call my "Exception Detection" as "Anomaly Detection".  In some cases the performance degradation could be caused by parasite program (like badly written data collection agent ) or incompetent user (like submitting badly written ad-hock  database query) or even by a cyber attack (denial-of-service attack -DoS definitely  degrades performance to absolutly not performing, doesn't it?)

So it is similar by my opinion and the Exception Detection methodology I am offering to by using MASF technique can be applied to broader filed of Information Security. And vice versa! Some intrusion detection techniques could be useful for automatic performance issues detection!

I have made a litle Google reserch on that and found a few interesting approaches. See one of that:

See the abstract page for dissertation written by Steven Gianvecchio:

Application of information theory and statistical learning to anomaly detection.


So the question is "can that information theory (entropy analysis) could be applied to performance exception detection?"

Friday, October 7, 2011

EV-Control Chart

I have introduced the EV meta-metric in 2001 as a measure of anomaly severity. EV stands for Exception Value and more explanation about that idea could be found here:  The Exception Value Concept to Measure Magnitude of Systems Behavior Anomalies 
Basically it is the difference (integral) between actual data and control limits. So far I have used EV data mostly to filter out real issues or for automatic hidden trend recognition. For instance, in my paper CMG’08 “Exception Based Modeling and Forecasting” I have plotted that metric using Excel to explain how it could be used for a new trend starting point recognition. Here is the picture from that paper where EV called “Extra Volume” and for the particular parent metric (CPU util.) it is named ExtraCPUtime:

The EV meta-metric first chart 

But just plotting that meta-metric and/or two their components (EV+ and EV-) over time gives a valuable picture of system behavior. If system is stable that chart should be boring showing near zero value all the time. So using that chart would be very easy (I believe even easier than in MASF Control Charts) to recognize unusual and statistically significant increase or decrease in actual data in very early stage (Early Warning!).

Here is the example of that EV-chart against the same sample data used in few previous posts:
1. Excel example: 

2.  BIRT/MySQL example as a continuation of the exercise from the previous post:

IT-Control chart vs. EV-Chart
Here is the BIRT screenshots that illustrate how that is built:

a.        A. Addition query to get EV calculated written directly in the additional BIRT Data Set object called “Data set for EV Chart”:
SQL query to calculate EV meta-metric
 SQL query to calculate EV metric from the data kept in MySQL table

B. Then additional bar-chart object is added to the report that is bind to that new “Data set for EV Chart”:
Result report is already shown here.





Tuesday, October 4, 2011

Building IT-Control Chart by BIRT against Data from the MySQL Database

This is just about another way to build an IT-Control chart assuming the raw data are in the real database like MySQL. In this case some SQL scripting is used.

1. The raw data is CPU hourly utilization and actually the same as in the previous posts: BIRT based Control Chart and One Example of BIRT Data Cubes Usage for Performance Data Analysis. (see the raw data picture here)

2. That raw data need to be uploaded to some table (CPUutil) in the MySQL schema (ServerMetric) by using the following script (sqlScriptToUploadCSVforSEDS.sql):

The uploaded data is seen at the bottom of the picture.

3.       Then the output (result) data (ActualVsHistoric table) is built using the following script (sqlScriptToControlChartforSEDS.sql):
The fragment of the result data are seen at the bottom of the picture also. Everything is ready for building IT-Control Chart and the data is actually the same as used in BIRT based Control Chart, so result should be the same also. Below is more detailed explanation how that was done.

4.  First, using BIRT the connection to MySQL database is established (to MySQLti  with schema  ServerMetrics to table ActualVsHistorical):

5. Then, the chart is developed the same way like that was done in BIRT based Control Chart post:


1.      6. Nice thing is in BIRT you can specify report parameters, that could be then a part of any constants including for filtering (to change a baseline or to provide server or metric names). Finally the report should be run to get the following result, which is almost identical with the one built for BIRT based Control Chart post: