Popular Post

_

Wednesday, June 30, 2010

IT-Control Chart against Network Traffic Data vs. Process Level Data Chart


The 1st implementation of applying SEDS methodology against network trafic data was published at my CMG'07 presentation (SYSTEM  MANAGEMENT  BY  EXCEPTION:
The Final Part):



Process level data chart could also be used together with the control charts to find which particular process is responsible for some unusual spikes. The figure above shows how the CPU usage by processes chart can be used to explain that incremental daily back-up causes small daily spikes on the control chart of Network traffic (NIC level). The full back-up caused one big spike per week expanding activity to work hours, which could be dangerous to interfere with other DB2 on-line workload on that server.



Monday, June 21, 2010

Industrial Robot Grasping Processes Research. EV prototype was there!


I have published  in my Russian blog here  the abstract of my  PhD dissertations  (1986 )

Research of Industrial Robot Grasping Processes 


The main idea of that work was to find the way to calculate a some set of initial  grasping (or assembling) object coordinates that would warrant successful grasping (or assembling) process (operation). I called that "Area of Normal Functioning - ANF" (Область Нормального Функционирования -ОНФ). If grasping or assembling process starts with parameters (or coordinates) that are not in that area (NFN), the process will be failed. That area defined coordinates where passive (natural) adaptation would work. Interesting that that robotic subject is still active - see the following link 


Underactuated hand with passive adaptation

Rereading that my old work I suddenly figured out that my resent idea of Exception Value (EV - area between statistical limits and just happened actual variables values) is very similar with that my very old idea of calculating limits for successful assembling or robot grasping processes!


Apparently my mind works very consistently.... 


   

Monday, June 7, 2010

Near-Real-Time IT-Control Chart R-Simulation

UPDATE: Now the following free web tool to build IT-control charts is available:
                           www.Perfomalist.com


See more explanation

Review of IT Control Chart

Wednesday, May 26, 2010

CMG'10 Interesting Paper Abstracts

CMG.org published the following interesting paper abstracts for the upcoming 2010 national conference that looks like are related to the subject of this blog:

    IT-Control Charts
    The Control Chart originally used in Mechanical Engineering has become one of the main Six Sigma tools to optimize business processes, and after some adjustments it is used in IT Capacity Management especially in “behavior learning” products. The paper answers the following questions. What is the Control Chart and how to read it? Where is the Control Chart used? Review of some performance tools that use it. Control chart types: MASF charts vs. SPC; IT-Control Chart for IT Application performance control. How to build a Control Chart using Excel for interactive analysis and R to do it automatically?

Effective Proactive Service Capacity Management using Adaptive Thresholds
Proactive Service Capacity Management (SCM) is essential to deliver desired service level to its users. It enables service to be available as per the SLA with agreed service performance. For proactive SCM, monitoring and proper thresholds are basic ingredients. Often service usage is seasonal in nature and setting fixed thresholds for alerts can’t take into account the variability of usage and significance of alerts. This paper is aimed at introducing the concept of adaptive thresholds and discusses how this should be utilized to proactively manage e2e service perf and capacity aspects. 
Using Statistical Process Control to Improve the Quality and Delivery of IT Services
This paper presents a framework for the delivery of IT services based on Continuous Quality Improvement (CQI). Starting with the Capability Maturity Model (CMM), we develop a process oriented approach based on Statistical Process Control (SPC). We apply the framework to the Change Management process of a large IT environment for a trading software firm, and show how failure-rates of the Change Management process were reduced dramatically. 


Thursday, May 20, 2010

Baselining and dynamic thresholds features in Fluke and Tivoli tools

I have just attended two demo sessions about Fluke VPM and Tivoli Monitoring tools and both presentations included the following new features related to performance Exception Detection technology:

1. Fluke (NetFlow Tracker - http://www.skomplekt.com/pdf/PerformanceAndScalability(fnet).pdf)
- Baselined alarms trigger when normal usage is exceeded.
- Automatically choose a threshold using a baseline, or manually specify.
- Tracker baselines individual elements of a report, not just the total.
- Baseline can be static (learn once) or update weekly (learn every week).

2. Tivoli Monitoring Version 6.2.2 ( http://publib.boulder.ibm.com/infocenter/tivihelp/v15r1/index.jsp?topic=/com.ibm.itm.doc_6.2.2/new_version622.htm)
- The bar chart, plot chart, and area chart have a new "Add Monitored Baseline" tool for selecting a situation to compare with the current sampling and anticipated values. The plot chart and area chart also have a new "Add Statistical Baseline" tool with statistical functions. In addition, the plot chart has a new Add Historical Baseline tool for comparing current samplings with a historical period.
- Situation overrides for dynamic thresholding.

That is interesting as I personally discussed with Tivoli team possibility to add SEDS like features to Tivoli tools just before I left IBM (about 3 years ago). It looks like they finally have implemented some elements of that technology!

I will be playing with all those features soon and plan to add more comment about how that works.

Tuesday, April 13, 2010

Disk Subsystem Capacity Management - my CMG'03 paper - "Health Index" metric and Dynamic Thresholds

Here is the link to my CMG'03 paper:  http://www.cmg.org/proceedings/2003/3099.pdf
(Free download but registration is required)
Presentation slides are freely available here:
Disk Subsystem Capacity Management, Based on Business ... - CMG

1. The paper showed interesting way to report Disk Space usage via BMC Perceive:


2. In the paper there is example of using some interesting "Health Index"  metric. I just took it from Concord (now it is CA product, I believe) performance data collector as one of many performance metrics.


Based on Concord eHeallth tool documentation:

“System Health Index” is the sum of five components (variables):
–SYSTEM, which reports a CPU imbalance problem;
–MEMORY, which is exceeding some memory utilization threshold or reflects some paging and/or swapping problems;
–CPU, which is exceeding some utilization threshold;
–COMM., which reports network errors or exceeding some network volume thresholds;
–And STORAGE, which might be a combination of
a. Exceeding user partition utilization threshold;

b. Exceeding system partition utilization threshold;

c. File cache miss rate, Allocation failures and

d. Disk I/O faults problem that can add additional points to this Health Index component.

I used that long ago. Currently in my environment I do not have that collector.
But I have started calculating my own way of "health index", which is based on numbers and types of exceptions (e.g. Hot ones are defects like run-aways; warning ones are just severe deviations from statistical norms; also number of hours/days with exceptions that does matter). Filtering that by applications (using CMDB) it gives you an idea of how stable the application is. In my other papers there are some elements of that approach.

2011 update: Other important  idea is in the paper is Dynamic Thresholds usage suggestion as for high level I/O related metrics there are no natural thresholds. Dynamic  Thresholds  got recently popular but I introduced that long ago!

Capturing Workload Pathology By SEDS - my CMG'05 paper

The paper can be found here: https://www.researchgate.net/publication/221447101_Capturing_Workload_Pathology_by_Statistical_Exception_Detection_System
Here is the resume:
Problem definition: The Servers workload pathology  (defects) such as run-away processes and memory leaks captures spare server resources and causes the following issues:
- being a parasite type of workload they compete for the resources with the real workload and causes performance degradations;
- they mimic capacity issue, but they are not a real capacity problem and just spoil the historical sample and causes wrong capacity trends as seen on the Figure below:

To fight this problem I have developed the way to capture those defects, report on them and then to remove them from historical sample to see real capacity trends. That was implemented as a part od SEDS application. Detailed explanations are in my CMG'05 paper  "Capturing Workload Pathology by Statistical Exception Detection System"
"Capturing_Workload_Pathology_by_Statistical_Exception_Detection_System)

Other good result of implementing this problem resolution was dramatic reduce number of incidents related to run-away and memory leaks defects. The chart below shows 2+ time reduction for 2 years:


Other work in this area made by Ron Kaminski. See CMG paper here:

Automating Process and Workload Pathology Detection


presentation slides:  

Automating Process and Workload Pathology

Tuesday, December 29, 2009

Exception Value (EV) and OPNET Panorama

I have recently looked at the following OPNET resources to get impressions of OPNET Panorama tool:

1. Link to website: http://www.opnet.com/solutions/application_performance/panorama.html
2. White paper downloaded from that site: "Understanding OPNET Panorama’s Performance Analysis Engines"

General comment: OPNET becomes the next generation tool provider that has “learning behavior” capabilities that are similar with what I do for years with my SEDS and with tools from other Vendors, such as Netuitive, Integrien and ProactiveNet (BMC) that I have recently studied (check my older postings: http://itrubin.blogspot.com/2009/02/realtime-statistical-exception.html).

Special comment: In the OPNET white paper I read: "Metrics that exhibit deviations from normal are automatically identified and assigned scores based on “how abnormal” their behavior is."
This is very close to what I introduced in my 1st CMG paper in 2001("Exception Detection System, Based on the Statistical Process Control Concept" ) and called ExtraVolume of a metric (I call that now Exception Value (EV) meta-metric). In OPNET’s white paper referenced to that my CMG’01 paper, but they did not mention that they use very similar approach to rang exceptions (“Area Out vs. Limit Range In Metric Scoring”).
Here is example from my 1st CMG paper of the usage EV metric to build the TOP exceptional Unix servers list:


I had even tied to normalize that metric to some Unix benchmarks (TPC) to compare ranges of exceptional capacity usage ecross diferent type of servers and configurations. The example of the report is in the paper.

Al in all, it is a good news that another vendor has been adopting that technology (maybe with some of my work’s influence!). Based on my experience with OPNET tools (very limited, just ITD Guru for network and some server behavior simulations), that tool most likely can be trusted. To speak more I need at least to play with demo….