Popular Post

_

Friday, August 31, 2012

Z Capacity Management without SAS and MXG

I have just commented the following post on LinkedIn:

"Is there anyone using anything else besides SAS or SAS assisted tools for analyzing Z/OS SMF data? If so, what are you using?"


I have involved in some IBM activity to offer an alternative (to SAS) solution to process and analyze SMF data by using TDS + SPSS + COGNOS (all IBM tools).

My part of this is to offer SETDS elements to include in the out-of-box COGNOS (and potentially SPSS) reportingsuch as IT Control Charts, EV based anomaly and recent trend detection.

Friday, August 17, 2012

How to Calculate Availability of Clustered Infrastructure for Multi-Tier Application

That is the task I am working on right now. I have some progress and the approach I found is to build availability graph to consider the clustered infrastructure as a chain of parallel and sires connected nodes described here with formulas. So below is a simple example:









And the availability calculation formula will be: 

A  = A1*(1-(1-(A2*A3)n)*A4

You can play with different level of redundancy "n" of the cluster here. Currently  it is 2 but you could estimate it for n=3 or n=4. That approach opens possibility to quantitatively justify you architectural decisions (not just using "best practices" or "gut feelings"). 

If you know MTTR for each individual component (SW and HW) you could estimate the whole infrastructure availability using this approach.  But how to get that individual MTTR? From vendors - good luck! Maybe from incident records? Or set up special monitoring for that (Synthetic- robotic?)

Other useful resources with formulas that relevant to this:

Thursday, August 2, 2012

SEDS-Lite: Using Open Source Tools (R, BIRT, MySQL) to Report and Analyze Performance Data - my new CMG'12 paper

20202 UPDATE: The SEDS-Lite web app is about to be released!
_________________________________________________________
I wrote this paper with some help from Shadi G. (from Dublin, also IBMer).
The paper is based on my blog postings:
SEDS-Lite Presentation at Southern CMG Meeting in the SAS Institute
SEDS-Lite Introduction
How To Build IT-Control Chart - Use the Excel Pivot Table!
BIRT based Control Chart

HERE IS THE VIDEO PRESENTATION
Below is the abstract:
Statistical Exception Detection (SEDS) is one of the variations of learning behavior based performance analysis methodology developed, implemented and published by Author. This paper took main SEDS tools – IT-Control Chart and Exceptions (Anomalies) Detector - and showed how that could be built by Open Source type of BI tools, such as R, BIRT and MySQL or just by spreadsheet. The paper includes source codes, tool screen-shots and report input/output examples to allow reader building/developing a light version of SEDS.
-------------------------
The presentation of this paper is scheduled on December 5th, 2012 Wednesday, 2:45:00 PM - 3:45:00 PM in Las Vegas, Nevada
-------------------------

THAT IS MY SECOND CMG'12 PAPER. THE FIRST ONE ANNOUNCED HERE:

AIX frame and LPAR level Capacity Planning. User Case for Online Banking Application

Tuesday, July 31, 2012

AIX frame and LPAR level Capacity Planning. User Case for Online Banking Application - my new CMG'12 paper

    I have just got acceptance notifications about my two new CMG papers I wrote and submitted for this year CMG'12 conference.
    Below is the abstract of the 1st one which is base on the successful project I had this year.
    AIX frame and LPAR level Capacity Planning. User Case for Online Banking Application
    The paper shares some challenges the Online Banking Capacity Management team had and overcame during the Solaris to AIX migration. The raw capacity estimation model was built to estimate AIX frames capacity needs. The Capacity planning process was adjusted to virtualized environment. The essential system, middleware and database metrics to monitor capacity were identified; business driver correlated forecast reports were built to proactively tune entitlements; IT-Control Charts were created to establish dynamic thresholds for Ph. Processors and IOs usage. Capacity Council was established. 

    The presentation of this paper is scheduled on December 5th, 2012 Wednesday, 9:15:00 AM - 10:15:00 AM in Las Vegas, Nevada (check updates here: http://www.cmg.org/conference/cmg2012/ )
    _____________________________________
    The 2nd paper information is on the next post:
    SEDS-Lite: Using Open Source Tools (R, BIRT, MySQL) to Report and Analyze Performance Data



Thursday, July 12, 2012

Just submitted CMG'12 papers abstracts: Very preliminary analysis

Abstracts are published anonymously here: http://www.cmg.org/cgi-bin/abstract_view.pl 
Apparently one of the  papers was inspired by me: 

Time-Series: Forecasting + Regression: “And” or “Or”?
At CMG’11, I had a fascinating discussion with Dr. I.Trubin. We talked about Uncertainty, Second Law of Thermodynamics, and other high matters in relation to IT. That discussion prompted this paper. We propose a method to get better predictions when we have a forecast of independent variable and a regression. It works for any scenarios where performance can be linked with business metrics. A real-world example is worked through that demonstrates how this technique works to improve the performance metric prediction and highlight trends that would have been overlooked otherwise.
 I guess that relates to my other posting about other paper that use "entropy" :  
Quantifying Imbalance in Computer Systems: CMG'11 Trip Report, Part 2
The following are abstracts of some other papers from the list that potentially could relate to the main topics of this blog. I cannot wait when I can read them!
Methods for Identifying Anomalous Server Behavior
Identifying anomalous server behavior in large server farms is often overlooked for a variety of reasons. The anomalous behavior does not breach alerting thresholds, or perhaps the behavior is subtle and is simply missed. Whatever the case, it is important to identify such behavior before it becomes more severe. In this paper we discuss methods of identifying server behavior that is anomalous or otherwise or uncharacteristic. Methods include statistical techniques such as multidimensional scaling, and machine learning methods such as isolation forests and self organizing maps.

Software Performance Antipatterns for Identifying and Correcting Performance Problems
Performance antipatterns document common software performance problems as well as their solutions. These problems are often introduced during the architectural or design phases of software development, but not detected until later in testing or deployment. Solutions usually require software changes as opposed to system tuning changes. This tutorial covers five performance antipatterns and gives examples to illustrate them. These antipatterns will help developers and performance engineers avoid common performance problems.


Introduction to Wavelets and their Application for Computer Performance Trend and Anomaly Detection
In this paper I will present a technique to identify trends and anomalies in Performance data using wavelets. I will answer the following questions: Why use Wavelets? What are Wavelets? How do I use them?

Application Invariants: Finding constants amidst all the change
This paper presents a method for deriving and utilizing Application Invariants. An Application Invariant is a metric that quantifies the behavior or performance of an application in such a way that its value is immune to changes in workload volume. Several sample Application Invariants are developed and presented. One of the primary benefits of an Application Invariant is that it provides a simple (flat) shape that can readily be used to track changes in application performance or behavior in an automated manner.
Couple other papers could be found there with the obvious interest for this blog.... Will post them later here.

All in all, based on the 1st glance, looks like this year CMG conference (http://www.cmg.org/ ) will have a great success.

Monday, July 2, 2012

Advanced process control (APC) and Fault detection and classification (FDC)

On my post "Virtual CMG'90 Trip Report about Control Chart UsageI have detailed and very interesting response from my 3rd LinkedIn connection Mike Clayton  from Engineering field (not IT at all!). That has a special interest for me as I came from that field originally (my 1st degree is in Engineering) and the SPC concept was originally designed for Engineering application and then adopted for IT via MASF in 1995.


Below is our dialog: 


MIKEUsing normally correlated parameters to detect "loss of correlation" as a fault, for example, is common now in monitoring process tools that have many sensors. Loss of expected correlation ties to actual physical faults, right? Is that one of the things you are finding in your history search? FDC as part of APC which has augmented SPC once we have found adjustment algorithms that can be automated based on output parameters IF the toolset or system passes the FDC check....otherwise, call for help?

____
I know nothing of IT performance metrics...except that most IT departments kow-tow to the Finance department, and not the operations dept. So this past year, our COO took over IT and we have been making great REAL performance progress since then at one of my clients.

But fault-detection is same everywhere I have found, in its multivariate nature, with attention to correlation structure changes.

Roy Maxion at CMU years ago wrote some code in old Xerox printer language (Postscript) that put out green-sheet graphs based on genetic algorithm looking at campus internet traffic.
It was amazingly effective for campus network support technicians.

I think Roy published in JOurnal of Machine Learning over the years. He loved the VAX OS...like me, but was very stubborn about doing anything on Windows OS for long time, so he missed the big money, but he was technically correct of course. I have great respect for Roy.

IBM's Ray Bunfkowski (spelling?) did pioneering work with APC methods at IBM semiconductor operations, and published for Sematech Workshops, and perhaps IEEE.

He often used Svante Wold's Umetrics software, an early pioneer of multivariate methods for engineers. Umetrics has a package called SimcaP I think.  


IGOR: Yes, it is! Looks like I intuitively went to the APC area applying some similar technique to Computer System performance data. Could you point me to any good books or paper about APC/FDC? Starting with basics... 


MIKE: http://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=05458323  
http://www-mtl.mit.edu/researchgroups/Metrology/PAPERS/goodlin-fault-detect-jecs2003.pdf 
http://www.umetrics.com/fabstat  
http://ieeexplore.ieee.org/xpl/login.jsp?tp=&arnumber=5398983&url=http%3A%2F%2Fieeexplore.ieee.org%2Fxpls%2Fabs_all.jsp%3Farnumber%3D5398983 

many more on web. Most interesting is to see how FDC and R2R work together now days in modern factories (the first reference above).

I first ran into this FDC issue at Motorola in 1990, and tried methods from vendors as well as universities. CMU's machine learning methods using Genetic Algorithms worked well for continuous processes, but were hard to use for discrete manufacturing where small bursts of data from one lot to the next had to be collected and compared based on start and stop signals without the batch run. Dumbing down from the slow learning but precise models of Neural Nets, to the faster learning and more robust models of Nearest Neighbors was part of my early learning.

Costas Spanos at Berkely, and Dr. Moyne at Univ of Michigan were big help since FDC ratings in realtime were needed to avoid over-adjusting from R2R feedback systems, interrupting to call engineering support before permitting tuning.

Sunday, June 3, 2012

Adrian Heald: A simple control chart using Captell Version 6

At the CMG'11 I have met Adrian and asked him to show me how his reporting tool “Captell” (www.reportingservices.com) can be used to build MASF Control Charts. Below is his response.Check also the comment to this post with my feedback.
_____________________________________________________________________________
Introduction


A control chart uses data from a specified period to derive average and upper and lower control values. For this example we are using some CPU utilization data from a UNIX machine collected over the 4 month period January through April 2011 and delivered in a CSV file. The baseline period is January, from which we calculate average values and standard deviations for each clock hour. We can then plot our control chart and compare successive month’s average with the control to see a clear picture of change.

For more information and sample reports see www.reportingservices.com
or contact
Adrian Heald
on +61 (0)411 238 755
adrian@reportingservices.com


Step 1 - Import the CPU utilization data.


The following dialog shows the table definition selecting the “Delimited text file” source type. Specify a name and folder and choose the source type.


Here we see the text file definition, all that is required is the filename and a specification of the date time format.


Step 2 – Import and view the data

This Window shows the main Captell dialog with the task importing the data


And here a view of the imported data; during the import of the data Captell automatically determines correct data types.

Step 3 – Create a query to calculate the base line

This query calculates the average and average +/- 2 standard deviations for data from January.


The query output.

Step 4 – Create a query to summarise single months data

This query calculates the average CPU for each hour throughout the month selected by the Captell parameter ‘Data\Month’.


The query output:

Step 5 – Create a chart to combine the two queries

This chart shows the baseline average CPU utilisation and upper control limit along with the average values from the current month. Captell’s ability to plot data from different sources, in this case the baseline data and the data from the new month makes reporting quite easy. The blue line with the square symbols shows the average hourly data for March, well within the control limit and all hourly values below the baseline average.


Step 6 – Change the parameter to compare a different month

Here we can see the parameter changed to April and the resultant chart. The blue line with the square symbols shows the average hourly data for April, mostly above the upper control limit and all but one hour above the January mean, indicating a substantial increase in utilization.




(Posted with the Adrian's Heald permission)




Wednesday, May 23, 2012

STEEDd: Another Implementation of The Near-Real-Time Control Charting and EV Calculating


Thierry Déléris is a French System Programmer on Mainframe in a team dedicated to performance, metrology & capacity planning. He used some ideas published in Trubin's CMG papers to implement the following:

1. The solution, wich gives a daily eMail by CEC with a spreadsheet by LPAR and Workload, on a daily basis: thresholds are calculated thanks to the R Language by day of the week, hour of the day, LPAR name and WLM Workload, based on a 6 month history data (based on SMF72 records) with exclusion of outliers using Tukey Statistical Method.
This initial part of the solution has a big inconvenient: it gives the resulting spreadsheet for a CEC only the next day because it is based on the SMF 72-3 records of the previous day collected during the last night by TDSz...

2. Then the second part of the solution called STEEDd (Statistical Tool for Enhanced Exceptions Detection and Diagnosis, and as a reference to the "Avenger" British TV Show character John Steed and is legendary bowl hat) was developed using a Java solution to use the same R calculated thresholds but on a 15 minutes control solution, which interacts with BMC Mainview on the Host to collect the current data (In fact the last 15 minutes data). This solution gives a main screen to select the metric to control, and a control screen by metric. An eMail alert is sent to the team if for some metric the result is higher or lower than the target high or low thresholds.

As an example, here is a picture of the control screen used for CPU Metric by Workload & LPAR :

Legend:
When the icon is selected, the associated control chart pops up showing the metric for the last 12 hours like the shown below:


The idea of EV (Extra Value or Exception Value, introduced in Trubin’s CMG papers and discussed in this blog) is used there (Red bars for EV+ and Yellow bars for EV- on above picture) . This helps filtering the right & false negative alerts.

3. Third part of the solution: On the way! An Artificial Intelligence solution based on a rule engine is studied to explore the detected problem by a hierarchical way... This application will be used to enhance the analysis of the metric alerts thanks to an "expert system" way.

(Posted with the Thierry Déléris permission)

Monday, May 7, 2012

SEDS-Lite Presentation at Southern CMG Meeting in the SAS Institute

Southern CMGLast Friday I have made my presentation which was announced here: SEDS-Lite: Using Open Source Tools (R, BIRT and MySQL) to Report and Analyze Performance Data. That was presented at the Southern CMG Meeting in the SAS Institute, Cary, NC. The presentation slides are linked within AGENDA and also can be downloaded from HERE


I plan to write a paper based on this presentation and to submit that to this year CMG'12 conference.


Friday, April 20, 2012

Building IT-Control Chart with COGNOS

I am developing SEDS elements using IBM Cognos. Here is the 1st result, which is just a POC prototype of IT-Control Chart report.
I used the test data (Date-hour stamped utilization metric) that I developed to build the same IT-Control Charts by other tools (BIRT, MySQL, R). I have published some information about that on my previous blog posts. (e.g. R-script to plot IT-Control Chart against MySQL)

This time I have developed simplest meta-data package against ODBC to MySQL database by using Cognos Framework Manager and published that in TCR locally on my Laptop. Then I used Cognos Report Studio to build the report. The result of running the report is following:

I got the same result as I got by using R or BIRT, but I have noticed some nice features in COGNOS that helped me to build that faster and more accurate (e.g. adding the dates at the X-Axis)

I am going to mention that progress with some details on my up-coming SCMG presentation:

SEDS-Lite: Using Open Source Tools (R, BIRT and MySQL) to Report and Analyze Performance Data

UPDATE: I will be presenting that again at CMG'12 conference:  http://itrubin.blogspot.com/2012/08/seds-lite-using-open-source-tools-r.html

Tuesday, April 17, 2012

Southern CMG Spring 2012 Meeting in Richmond - MXG is our Sponsor!

At SCMG we have usually two meetings each season (2 Spring and 2 Fall ones, both in Richmond and Raleigh). Last season - 2011 fall - I had my presentation; see the following post: "My Southern CMG Presentation in Richmond Is About Open Source Tools for Capacity Management ". Presentation slides are published here: slides

This spring I have the similar but updated presentation in our Raleigh SCMG meeting: "SEDS-Lite: Using Open Source Tools (R, BIRT and MySQL) to Report and Analyze Performance Data" 

So this time I am not presenting in Richmond but I found very good sponsor for that meeting - Merrill Consultants (http://www.mxg.com). Barry Merrill himself responded on my invitation and now we have a great opportunity to see and listen the legendary Capacity Management inventor!

Please consider attending our Richmond VA SCMG meeting on May 11, 2012: 
 
http://regions.cmg.org/regions/scmg/spring_12/richmond/meeting.htm


Wednesday, April 11, 2012

SEDS-Lite: Using Open Source Tools (R, BIRT and MySQL) to Report and Analyze Performance Data

My presentation with this name has been scheduled for the next Southern CMG meeting  at SAS Institute:

SCMG Meeting Raleigh
May 04, 2012

You are welcome to attend!

Thursday, April 5, 2012

Prehistory of SEDS: Virtual CMG'90 Trip Report about Control Chart Usage. Part 1.

Using the key word "Control Chart" I have found in the www.CMG.org knowledge base a few very old CMG papers with some discussions about using classical SPC approach against computer performance data.

Here is the first one:

 Fine-Grain Analysis (FGA): A Methodology for Analyzing Intermittent Performance Problems Open in a new window
  By Robert Berry & Jeffrey Hedglin 

 

The paper describes what Mainframe metrics are good to use for Control Charting. They should be two types - a. Performance Quality Measure - sounds like modern KPI... (e.g. response time);  b. System performance metrics (e.g. CPU queue length). Then the paper describes how the intermittent problem could be detected just by plotting SPC Control Charts for both type of metrics in sync (correlated).

I use that approach a lot now, but using MASF type of Control chart and specifically my IT-Control Charts.  BTW I am writing now my next CMG paper and plan to add there a couple very persuasive  examples of correlated IT-Control Charts, such as, number of concurrent user LOGONS vs. number of Ph. CPUs used by LPARS on some p770 AIX frame....

To be continued....