Ankit Kulkarni
Members
-
Joined
-
Last visited
Solutions
-
Ankit Kulkarni's post in Fix Fast or Fix Right โ What Should AI Drive? was marked as the answerMy position
I donโt agree with prioritizing learning over immediate resolution.
In real operations, the sequence is critical:
๐ Contain first (fix fast)
๐ Then eliminate root cause (fix right)
Reversing that order increases operational exposure, cost of poor quality (COPQ), and system instability.
Example 1 โ Power plant operations (real-time incident response)
In our power plant operations, predictive analytics continuously monitors turbine health โ vibration, temperature, load fluctuations โ to detect early failure signals.
When a turbine trips or shows abnormal behavior, the impact is immediate:
Drop in available capacity (MW loss)
Reduced plant availability and load factor
Increased risk of forced outage
Direct revenue loss per hour
At that point, the system is outside stable operating conditions โ effectively beyond control limits.
LSS framing
Containment action โ restore the process within control limits (bring unit back safely)
Corrective action โ identify and remove assignable cause
Preventive action โ modify control strategy to improve MTBF and reduce recurrence
If we delay containment in favor of analysis:
Availability loss increases
Throughput drops
Risk of cascading failures rises
COPQ escalates rapidly
So in practice:
๐ We stabilize first (containment)
๐ Then run structured RCA (5 Why, fishbone, failure mode validation)
๐ Then strengthen controls (FMEA updates, predictive thresholds, SOP changes)
Example 2 โ LCD carrier rejection (manufacturing case)
In a previous role, we encountered severe distortion in LCD carriers post injection molding:
Rejection rate reached ~85%
Production line supporting ~โฌ85M revenue was at risk
Effective throughput collapsed
At that point, the process capability had clearly shifted โ a classic special cause variation scenario.
Step 1 โ Containment (fix fast)
Introduced rework process
Achieved ~35% recovery rate
Maintained partial line output
From an LSS lens, this was:
๐ Containment to reduce immediate COPQ and throughput loss
Step 2 โ Corrective & Preventive (fix right)
We then moved into structured DMAIC:
Measure โ MSA at supplier and plant
Analyze โ cycle time (7.5 min), cooling fixture (10 min), material behavior
Root cause โ deformation due to cooling fixture design
Improve โ mold and process redesign
Control โ updated specs, monitoring limits, supplier controls
๐ Result: Rejection reduced to near zero, process returned within stable limits
What this shows
If we had focused only on learning first:
Production would have stopped completely
Availability and throughput would collapse
COPQ and revenue loss would escalate
If we had focused only on quick fixes:
Recurrence probability remains high
MTBF remains low
System stays in firefighting mode
The correct sequence is:
๐ Contain โ Correct โ Prevent
Why Bexโs position is incomplete
I agree that root cause elimination is essential.
But prioritizing it before stabilization ignores real-world system dynamics.
Because:
RCA requires stable conditions
Data collected during instability is often misleading
Extended disruption increases operational risk and cost
In LSS terms:
๐ You cannot run a reliable Analyze phase when the process is not under control
What AI should actually drive
AI should not force a trade-off between speed and learning.
It should enhance the full improvement cycle:
Faster detection โ earlier containment
Pattern recognition โ sharper root cause hypotheses
Feedback loops โ stronger preventive controls
AI improves both reaction speed and learning depth โ
but the sequence must remain disciplined.
Bottom line (my view)
From a Lean Six Sigma perspective:
๐ Containment protects availability, throughput, and customer impact today
๐ Corrective and preventive actions improve capability, reliability, and MTBF tomorrow
AI should accelerate both โ
but never confuse their order.
-
Ankit Kulkarni's post in Should AI Stop the Process Before a Defect Happens? was marked as the answerMy position
I donโt support stopping the process every time AI raises a flag.
In my view, reacting to every prediction โ even at 85โ90% accuracy โ creates a different problem:
we reduce defect risk, but we damage flow, availability, and trust in the system.
Example from operations (Power plant โ turbine reliability & outage strategy)
Let me take a real example from our operations, we have 54 GW power generation coming from CCGTs under different locations.
We use predictive analytics on our turbines data โ vibration, temperature, load behavior โ to flag potential failures early.
Now, in this environment, decisions are not just โstop or continue.โ
They are tied to outage strategy:
Minor outage โ few hours to 1โ2 days
Major outage โ several days to weeks
Both decisions have cost:
If we miss a real failure, we face forced outages, generation loss, and possible equipment damage
If we shut down unnecessarily, we lose generation revenue and disrupt grid commitments
So we donโt treat every signal the same.
We use a risk-tiered response, not a binary stop rule.
What a risk-tiered response looks like
๐ด High-risk signals
Strong deviation across multiple parameters
Matches known failure patterns
High impact on critical equipment
๐ Action: Immediate controlled shutdown
(Shift into planned outage โ minor or major depending on severity)
๐ก Medium-risk signals
Moderate deviation
Early-stage anomaly
๐ Action:
Continue running
Increase inspection / monitoring
Prepare for planned intervention
๐ข Low-risk signals
Weak or inconsistent signals
๐ Action:
Monitor only
No disruption
Bringing it back to the given scenario
Letโs look at the numbers:
3โ4 flags per shift
12% false positives
15โ40 minutes per stoppage
If we stop on every flag:
Total stoppage per shift
45 to 160 minutes per shift
On an 8-hour shift (480 minutes):
๐ Thatโs ~9% to 33% loss of availability
False positives alone
0.36 to 0.48 unnecessary stops per shift
Equivalent to ~5 to 19 minutes lost per shift
๐ Thatโs ~1% to 4% pure false-alarm loss, before even considering restart instability or downstream effects.
Daily impact (3 shifts)
2.25 to 8 hours lost per day
Of which ~16 to 58 minutes is pure false positive loss
At that point, we are no longer protecting quality โ
we are systematically disrupting flow.
The real lens: Type I vs Type II error
This is fundamentally a Type I vs Type II trade-off:
Type I error (false positive) โ stopping when no defect would occur
Type II error (missed defect) โ not stopping when failure is real
Stopping on every signal tries to eliminate Type II error โ
but at the cost of very high Type I error impact.
In operations, both errors have cost.
The goal is not to eliminate one โ
it is to balance them intelligently.
Where I differ from Bex
I agree with Bex on one point:
defects reaching the customer are unacceptable.
But stopping every time AI predicts risk is not discipline โ
it is treating probability as certainty.
AI gives early signals.
Operations must convert those signals into proportionate action.
Bottom line (my view)
AI should trigger attention, not automatic interruption.
Stopping the process should be reserved for:
๐ high-confidence signals + high-impact risk
Everything else should be handled through:
Monitoring
Containment
Planned intervention
Otherwise, we solve one problem โ defects โ
by creating another: loss of flow, trust, and operational discipline.
-
Ankit Kulkarni's post in When Should People Trust an AIโs Recommendation โ and When Should They Override It? was marked as the answerProcess Context
My team also manages the central master data management for 50+ plants today, and this will grow to 70+ plants by 2027. The entire fleet data management is handled by two people in my team.
Every time we commission a new plant or acquire a new plant, we need to align itโs material master with our fleet database, to avoid duplication, planning errors, and wrong spares being introduced into SAP.
In a typical post commissioning and acquisition, we review minimum of 5000+ incoming material records against an existing 345000 item fleet master.
Practically for my team, each item review takes at least 6 minutes without AI.
The AI-enabled Process
To solve this, I built a Python + AI solution using a MiniLM semantic model, combined with rule based checks.
The program setup classifies each incoming item into three categories, Auto, high confidence match to directly map & upload in SAP. Review, ambiguous match, reviewed by the master data team. Reject, no valid match, program generates a new master data creation template for my team, to directly load into SAP.
You can clearly see, AI does not create master data blindly in this case, it recommends, and the team decides.
When We Trust The AI
I have defined clear rules after testing the model for almost 10 days with millions of lines, semantic similarity is high & critical identifiers (model number, size, rating). It checks if descriptions and attributes are complete and consistent. One more rule I have setup is to keep standard, low-risk categories, and excluding verified MRP items, and these items directly flow straight into Auto category & are uploaded without manual touch.
When We Override The AI
Team deliberately does the review when similarity scores are close across multiple candidates, technical digits conflict even if text similarity is high. Then we also look at if item is maintenance critical or safety critical. We jump to the poor descriptions as well.
In all such cases, teamโs priority is correctness, not the speed.
ย
Safeguards That Keep The Balance
We have built simple controls to avoid blind trust or even excessive overrides,
Strict thresholds for Auto classification, mandatory teamโs review for all Review cases, spot audits of Auto mappings, tracking & analysis of override patterns to improve program, and we have clear ownership, AI suggests, Team decides.
Impact In Real Numbers
Now with this program, my team completes 5000 item migration in 10 days in total instead of 2 months.
I have a clear breakdown of 10 days,
Data setup + AI pre-load + first analysis is done in 0.5 day
SAP mapping for Auto category takes 1 day
Manual review is done for Review category in 7 days
New MD setup for Reject category is done in 1.5 days
This has really improved my teamโs output and bandwidth, and also reduced the onboarding risk for new plants, and best part is, it is allowing two people to scale this work for our growing fleet.
Bottom Line
I trust AI where signals are strong & mistakes are low impact, I override it where ambiguity or risk is high. As you can see, we are improving the overall process, idea isnโt to remove people from the process, itโs to make sure people spend time only where judgement actually matters.