
Gideon Marsh
The format problem comes before the search problem, and a great deal of time goes into searching logs that were never going to answer the question. Unstructured lines mean a pattern, and the pattern is a regular expression over the whole line, which is slower and more fragile than it looks. Timestamps are the recurring failure. ISO 8601 sorts lexicographically and works. A local time without a zone does not sort correctly across a daylight saving change, and the same log file can contain two formats because two services produced it. Normalising at write time is far cheaper than normalising at read time. Identifiers are the other one. A request that crosses three services needs a correlation identifier carried through all three, and without it the only way to follow a request is to match on timing, which fails under load. Redaction at write time is the only reliable redaction. Anything written to a log has usually been stored somewhere you do not control, so removing it later is a request rather than a fix. Volume drives the tooling choice, and I cover which questions need which approach: the recent and the occasional, against months of data across many services. Retention is the decision that constrains all the others, and it is worth setting deliberately: a shorter window makes a log cheap and makes an investigation that starts late impossible. Sampling is the compromise most teams settle on: keep everything for a short window and a fraction of it after that, which preserves shape and loses detail.
About ToolSura
ToolSura offers 80+ free, privacy-first online tools that run 100% in your browser — no uploads, no logins. Learn more about our mission →