On my post "Virtual CMG'90 Trip Report about Control Chart Usage" I have detailed and very interesting response from my 3rd LinkedIn connection Mike Clayton from Engineering field (not IT at all!). That has a special interest for me as I came from that field originally (my 1st degree is in Engineering) and the SPC concept was originally designed for Engineering application and then adopted for IT via MASF in 1995.
Below is our dialog:
MIKE: Using normally correlated parameters to detect "loss of correlation" as a fault, for example, is common now in monitoring process tools that have many sensors. Loss of expected correlation ties to actual physical faults, right? Is that one of the things you are finding in your history search? FDC as part of APC which has augmented SPC once we have found adjustment algorithms that can be automated based on output parameters IF the toolset or system passes the FDC check....otherwise, call for help?
____
I know nothing of IT performance metrics...except that most IT departments kow-tow to the Finance department, and not the operations dept. So this past year, our COO took over IT and we have been making great REAL performance progress since then at one of my clients.
But fault-detection is same everywhere I have found, in its multivariate nature, with attention to correlation structure changes.
Roy Maxion at CMU years ago wrote some code in old Xerox printer language (Postscript) that put out green-sheet graphs based on genetic algorithm looking at campus internet traffic.
It was amazingly effective for campus network support technicians.
I think Roy published in JOurnal of Machine Learning over the years. He loved the VAX OS...like me, but was very stubborn about doing anything on Windows OS for long time, so he missed the big money, but he was technically correct of course. I have great respect for Roy.
IBM's Ray Bunfkowski (spelling?) did pioneering work with APC methods at IBM semiconductor operations, and published for Sematech Workshops, and perhaps IEEE.
He often used Svante Wold's Umetrics software, an early pioneer of multivariate methods for engineers. Umetrics has a package called SimcaP I think.
IGOR: Yes, it is! Looks like I intuitively went to the APC area applying some similar technique to Computer System performance data. Could you point me to any good books or paper about APC/FDC? Starting with basics...
MIKE: http://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=05458323
http://www-mtl.mit.edu/researchgroups/Metrology/PAPERS/goodlin-fault-detect-jecs2003.pdf
http://www.umetrics.com/fabstat
http://ieeexplore.ieee.org/xpl/login.jsp?tp=&arnumber=5398983&url=http%3A%2F%2Fieeexplore.ieee.org%2Fxpls%2Fabs_all.jsp%3Farnumber%3D5398983
many more on web. Most interesting is to see how FDC and R2R work together now days in modern factories (the first reference above).
I first ran into this FDC issue at Motorola in 1990, and tried methods from vendors as well as universities. CMU's machine learning methods using Genetic Algorithms worked well for continuous processes, but were hard to use for discrete manufacturing where small bursts of data from one lot to the next had to be collected and compared based on start and stop signals without the batch run. Dumbing down from the slow learning but precise models of Neural Nets, to the faster learning and more robust models of Nearest Neighbors was part of my early learning.
Costas Spanos at Berkely, and Dr. Moyne at Univ of Michigan were big help since FDC ratings in realtime were needed to avoid over-adjusting from R2R feedback systems, interrupting to call engineering support before permitting tuning.
This blog relates to experiences in the Systems Capacity and Availability areas, focusing on statistical filtering and pattern recognition and BI analysis and reporting techniques (SPC, APC, MASF, 6-SIGMA, SEDS/SETDS and other)
Popular Post
-
I have got the comment on my previous post “ BIRT based Control Chart “ with questions about how actually in BIRT the data are prepared for ...
-
Re-posting interesting article from R-vs.-Python-for-Data-Science R vs. Python for Data Science Norm Matloff, Prof. of Computer ...
Monday, July 2, 2012
Advanced process control (APC) and Fault detection and classification (FDC)
Igor Trubin began his engineering career in 1979 as an IBM/370 systems engineer. He earned a Ph.D. in Robotics from St. Petersburg Technical University in 1986 and spent 12 years there as a professor teaching CAD/CAM and Robotics. He has published and presented more than 60 technical papers and conference presentations in robotics, artificial intelligence, IT performance, capacity management, anomaly detection, and FinOps. After moving to the U.S. in 1999, Igor worked at Capital One, IBM, and SunTrust Bank in senior engineering, architecture, and management roles. His 2002 CMG paper on exception detection received a Best Paper Award. He later developed Perfomalist.com, based on his original methods for anomaly, change-point, and trend detection, and created the online course “Performance Anomaly Detection.” At Capital One, he led development of cloud capacity-management and FinOps solutions, including the award-winning OptiCloud application. He has served on the CMG Board of Directors since 2015. Now semi-retired, Igor focuses on research, writing, and consulting. His current work expands his concept of the “Area of Normal Functioning” (ANF) from technical systems to human, orga
2 comments:
BB0A1E367C
kiralık hacker
hacker arıyorum
kiralık hacker
hacker arıyorum
belek
F6997393
Kütahya
Gümüşhane
Edirne
Aksaray
Balıkesir
Kırklareli
Muş
Ardahan
Urfa
Post a Comment