Find the real reason the line keeps stopping
Industrial troubleshooting and optimization is engineering applied to the problems that linger: the intermittent fault no one can catch in the act, the line that runs 20% below nameplate for reasons everyone explains differently, the scrap rate that crept up and stayed. It removes the cycle of guessing, part-swapping, and living with it.
Willowark troubleshoots with instrumentation instead of opinion. We put measurement on the problem — high-speed logging on PLC tags, current and vibration sensors, power quality monitoring, cameras on the failure point — and let the data identify the mechanism. Intermittent problems stop being intermittent when something is watching every occurrence.
Industrial AutomationHow the work gets done
The same way every time: scope, build, hand over.
The method is systematic: define the fault precisely — what, when, under what conditions — instrument to capture it, then analyze, correlating the event against everything logged in the seconds before it. Tools range from PLC trace captures at scan-time resolution and oscilloscopes on suspect signals to thermal imaging, airflow and pressure measurement, and statistical analysis when the problem is drift rather than an event. Optimization follows the same discipline: measure the current state honestly, identify the constraint, change one variable at a time, and verify with data — designed experiments where interactions are suspected.
The engagement ends with the mechanism named and fixed, not just the symptom suppressed — and intermittent culprits are usually mundane once caught: the kind of ground fault that appears only when a nearby crane runs, a photo eye washed out by afternoon sun, one label printer drifting out of spec. We document root cause, fix, and evidence, and where the fix needs monitoring to hold, we leave instrumentation in place watching for recurrence.
An engagement starts with a site walk and a careful interview: who has seen the fault, what was happening at the time, what has been tried, and what the operators notice that nobody wrote down. We read the fault history from the controller and drives, check the electrical basics — grounding, power quality, loose terminations — and review the last changes to the machine and its program. That produces a short list of candidate mechanisms and a measurement plan for each: which signals to log, at what resolution, and what condition should trigger a capture. Instrumentation is installed to that plan, usually during a planned stop.
Temporary instrumentation is installed without altering safety circuits or defeating guards, and any test that requires running outside normal conditions is planned with your safety and maintenance staff in advance. The usual traps are known and designed against: a capture window too short to see the cause, a suspect fixed prematurely so the evidence disappears, and a correlation mistaken for a mechanism. We log broadly enough to see the seconds before each event and change one thing at a time. Closure includes a written root cause report with the data, the corrective action and its verification, and any monitoring left behind, so the finding survives staff turnover.
Scope it in writing
What we agree before work starts
- Problem definition with fault conditions and history
- Temporary instrumentation to capture the failure in the act
Build with checkpoints
Working results, not slide decks
- Root cause analysis with supporting data
- Implemented fix with verification measurements
Hand over something you own
Documentation, source, and training
- Recurrence monitoring where warranted
- Written root cause report with captured data, corrective action, and verification results
Sound familiar?
Where industrial troubleshooting & optimization earns its keep.
A servo fault that appears twice a week, never during a service visit
A line running 20% under nameplate with five competing theories why
Scrap that doubled after a material change nobody can prove is the cause
A machine the OEM, the integrator, and maintenance have all blamed on each other
Ask about Industrial Troubleshooting & Optimization
Describe the problem. Get a straight answer.
One line is enough. An engineer replies within a business day.
Related work
Radar centering system for steel mills
Components:
- Radar sensors (strip position): Non-contact radar reading strip edge position in a hot, dusty, vibrating environment where optical sensors fail.
- Edge controller (signal processing): Turns raw radar returns into a clean lateral offset in real time.
- Mill PLC (centering actuators): The mill's existing controller: receives the offset and drives the centering actuators.
- Operator HMI (live position): Live strip position for the operator.
Connections:
- Radar sensors to Edge controller (raw returns)
- Edge controller to Mill PLC (offset)
- Edge controller to Operator HMI
A steel-mill systems provider · Steel manufacturing
Radar-based centering system for steel mills
End-to-end engineering of a radar sensing system that measures and centers material on steel mill lines — from equipment assessment through hardware selection, electrical engineering, software, installation, and commissioning.
Read the case study →Common questions
Asked before every industrial troubleshooting & optimization project.
We have already had three people look at this. Why would you find it?
Because we instrument rather than inspect. Most stubborn problems have been examined while behaving normally; the fix is watching continuously with high-resolution capture so the failure is recorded the moment it happens. Data catches what walkthroughs cannot.
Can you troubleshoot without stopping production?
Usually, yes — most instrumentation installs during planned stops or in minutes, then observes while you run. Running is the point: the problem happens during production, so that is when measurement matters. Interventions that require downtime get scheduled with you.
What if the root cause turns out to be a design problem, not a fault?
Then you know it with evidence, and you have options costed against each other — a modification, an operating change, or an engineered workaround. An answer backed by data that the machine cannot do what it is being asked to do still ends the guessing.
How long does it take to catch an intermittent fault?
It depends on how often the fault occurs and whether the first instrumentation plan is watching the right things. A problem that appears a few times a week is usually captured within a few weeks of monitoring; rarer faults take longer, and we occasionally need to widen or move the instrumentation after the first capture reveals a surprise. We report progress as we go rather than waiting for a full answer.
Can you also help us get more out of a line that is running fine?
Yes — optimization is the same discipline pointed at a different question. We measure the current state honestly, identify the constraint, and test changes one variable at a time with data confirming each gain. Typical targets are cycle time, changeover duration, scrap, and energy use. The important honesty is that not every line has hidden capacity, and we will tell you when the measured constraint is a physical limit rather than a tuning opportunity.
Where this sits
Industrial Troubleshooting & Optimization, inside a industrial automation system.
The lit component is the part of the system this service delivers; the rest is what it has to work with.
Hover or focus a component to see what it is and what it talks to. Arrow keys move between them.
Sensors and machines report into the PLC; the PLC drives the HMI and publishes tags to a historian; the historian feeds dashboards and, where it exists, the MES.
Components:
- Machine (press, cell, line): The equipment itself. Newer machines expose tags; older ones need a sensor or a serial tap.
- Sensors (counts, temps, current): Retrofit sensors where the machine offers nothing: proximity counts, current transformers, temperature.
- PLC (control): The controller: logic, safety, and the tag table everything else reads.
- HMI (operator): Operator screen at the machine.
- Historian (tag store): Time-series store of PLC tags: uptime, counts, faults, cycle times.
- Dashboards (OEE, downtime): Plant TV and office views: OEE, downtime reasons, shift comparison.
- MES / ERP (orders): Work orders down, production counts up.
Connections:
- Sensors to PLC over 4-20mA
- Machine to PLC over EtherNet/IP
- PLC to HMI over EtherNet/IP
- PLC to Historian over OPC-UA
- Historian to Dashboards over REST
- Historian to MES / ERP over REST, both directions
Strategy. Software. Systems.
Have a system that should exist?
Tell us what your operation is doing manually, what isn't connected, or what you're trying to build. We'll tell you plainly whether and how we can help.
