The process was running. The system was dead.
A couple of weeks ago, one of the systems we built and support stopped doing its job for about 42 hours, and almost nothing looked wrong while it happened.
It’s a document-processing system: documents arrive by email; an AI step reads the PDFs and turns them into structured data; our application validates and processes them; and the results flow out to the customer’s system of record. They’d not long gone live and were steadily onboarding new clients, so the volume was climbing week on week. Over a weekend, the scheduled processing simply stopped. The web application stayed up; the Java process never fell over; and every check watching it kept reporting the service as healthy. No crash, no alert. It only surfaced because someone noticed documents weren’t coming through, not because anything told us.