Build a durable event writer
Many threads, one log, and a promise that append returns only once the event is on the disk: positions handed out under a lock, callers who arrive together sharing one sync, and a reader that stops at the last whole record.
What is group commit, and why does it help?
A sync is the slow call, and a design with one sync per event pays it once for every caller. Group commit lets every caller that arrives while a sync is running wait for the next one together, so one sync covers all of them. The caller that finds nobody syncing takes the whole pending batch, lets go of the lock, writes and syncs once, and wakes everyone in it. In this build the same eight events under the same schedule cost eight syncs with a lock around each and four with group commit, and nobody is acknowledged any earlier than their own event is durable.
How do you make file order match position order?
Hand out the position under the same lock that adds the event to the pending batch, so the batch is always in position order. Then let only one caller write at a time, and have it take the whole batch in one step. Batch after batch lands in the order the positions were given, and a test can check it by sorting the acknowledgements by position and comparing them with the file. Positions handed out anywhere else, or two writers allowed at once, and the promise is gone.
Does write() returning mean the event is on the disk?
No. A write into a buffered Python file lands in Python's own buffer; flush() forces those bytes into the raw stream, which is the operating system; and os.fsync() is what forces the file's data to the disk. The Python documentation says to do flush() first and then os.fsync(f.fileno()). A writer that acknowledges after write() but before the sync has told its caller something a power cut can make false, and this build records exactly that happening.
What should the writer do when a sync fails?
Fail every caller whose event was in that batch with an exception that names the failure, and chain the original error to it. Then decide what later appends get. This build refuses them: after a failed sync nobody can say which of the batch's bytes reached the device, and handing out further positions would put events after a hole of unknown size. A caller that got the exception has to treat its event as unknown rather than lost, because some of the bytes may have landed.
How does recovery cope with a write torn by a crash?
Frame every record with its length and a checksum of its body. Recovery reads a header, checks that the whole body is there and that its checksum matches, and stops at the first record that fails either test. A torn tail is then dropped rather than misread, and a record that was never acknowledged may still be read back if all its bytes happened to land, which is allowed: the promise is that every acknowledged event is in the file, not the other way round.