Nazeem.Me

A blog about technology, football and all the other random stuff in my life

The Detector Was Never Looking

Written by

in

·

A neural network running at 5.8 milliseconds an inference, pointed at a room with cats in it, that had never once detected a cat.

I have an NVR. It runs in a container on one of the cluster nodes, it does object detection on that node’s integrated GPU, and it is fast: 5.84 milliseconds per inference, sitting at about 6% of what the detector can do. Person events come in with confidence scores between 0.73 and 0.93, which is healthy.

Cats walked past it for days. It recorded nothing.

The reflex is obvious and I had it immediately. The model is ssdlite_mobilenet_v2 at 300×300 — the entry-level detector, the one everyone upgrades away from. Every forum thread on the subject ends with someone recommending a dedicated inference accelerator. I had a tab open on where to buy one.

Then I opened the config.

objects:
  track:
    - person

That was the whole list.

The model could already do it

cat is class 18 in the label file the running model was loaded with. dog is 19. Both had been there the entire time.

The track list is not a hint passed into the network. It is a filter applied to what comes back out. The detector had been finding cats perfectly well and the NVR had been discarding every one of them before they reached an event.

So the thing I was about to spend money to speed up was work that was already being thrown away.

Diagram of a detection pipeline: the detector finds both a person and a cat, then a track list filter set to person only discards the cat before it can become an event.
The detector found the cat every time. The track list discarded it before it could become an event.

What the second class actually cost

BeforeAfter
Inference5.84 ms6.1 ms
Detections per second5.110.6
Skipped frames0.00.0

Detections per second roughly doubles, which is exactly what you would expect from going from one class to two. Nothing is skipped. The detector is still at around 6.5% of capacity. The second class is, in any sense that matters, free.

I had been treating the track list as if adding to it were expensive. It is not expensive. It was never the constraint.

Cats are not small people

The one genuine subtlety is thresholds. The defaults for a person on this model are 0.7 confidence and a 0.5 minimum score. I set cats to 0.6 and 0.4.

Cats score lower on mobilenet-class detectors. They are smaller in frame, they deform far more than a standing human, and they spend most of their time partly behind furniture. Start them at person values and you drop real detections — which is the normal route to concluding “it can’t detect cats” about a model that can.

That cuts both ways, so the thresholds are provisional until there is evidence. Person events on this camera land at 0.73 to 0.93. If cat events come in at 0.85 and above, the thresholds are too loose and I will tighten them toward the person values. If they come in low, 0.6/0.4 is what caught them at all. Raise the minimum score before the confidence threshold.

One habit worth keeping: validate the config file before restarting. The NVR has a validate flag that parses the config and tells you it is good. An indentation error you find that way costs nothing. The same error found by restarting takes the recorder down.

The accelerator I didn’t buy

Before ordering hardware I checked what each detector backend in the running container actually supports.

BackendModel families
Current, on the integrated GPUdfine, rfdetr, ssd, yolonas, yologeneric, yolox
The USB accelerator everyone recommendsssd, yologeneric

The accelerator removes four of the six families, including the two most accurate ones, in exchange for inference no faster than the iGPU already delivers. It would have been a downgrade I paid for, installed, and then written a post about being pleased with.

The published figures make the same point. The documentation quotes roughly 15 ms for this model on this class of integrated GPU. This box measures 5.84. So treat the numbers below as a pessimistic ceiling.

ModelResolutionInferenceCameras at 5 fps
The current one3005.84 ms measured34
One tier up320~16 ms12
Two tiers up320~20 ms10
Top accuracy tier640~46 ms4

Even the top tier at full resolution leaves this single camera at about 23% of the detector. There is headroom for a better model. There is no case for different hardware.

Two honest caveats if you go that route. There is no downloader for object detection models — the model cache directory is empty and stays empty, and a newer model has to be exported to ONNX by hand, with a config block whose tensor layout and data type differ from the current one. Getting those two fields wrong is the usual failure. And the “no contention between hardware transcoding and detection” result I recorded for this node was measured at 5.84 ms. At 46 ms the GPU is doing eight times the detection work on a chip that is also transcoding video. That finding does not carry over automatically and I would re-measure it.

Bar chart of inference time per detection model, showing 5.84 milliseconds measured against roughly 15 milliseconds documented, and the number of cameras each model supports at 5 fps.
Measured inference against the published figures, with the cameras each model supports at 5 fps. One camera is in use.

Why it stayed invisible for so long

This is the part that generalises past NVRs.

journalctl -u <service> on this thing shows systemd’s own start and stop lines and nothing else. The unit redirects both output streams to a file on tmpfs. So the application logs live somewhere the journal has never heard of, and they are wiped on every restart.

Which means: grepping the journal always comes back clean. Always. That cleanliness is not evidence of anything, and I spent time reading it as if it were.

It has a second consequence I ran into the hard way. The admin password is generated once on first run and printed to that log. Restart the service and it is gone — not rotated, just unreadable. Recovering it means setting a reset flag in the config, restarting, reading the new password out of tmpfs before anything else restarts the service, and then removing the flag again, because left in place it issues a new password on every boot.

The failure that looked like hardware and wasn’t

One more, because it has the same shape.

Early in the build every detector plugin died on startup with an exception out of a C++ standard library call. The obvious reading was a GPU passthrough problem — the container is unprivileged, the render device is bind-mounted in, that is where the risk lives.

It was the core count. The container was capped at 6 cores on a 12-thread host. Container filesystems virtualise the CPU list in /proc/cpuinfo down to the cap, but they do not filter /sys/devices/system/cpu/, which still exposes all twelve. The inference runtime walks sysfs, looks each CPU up in cpuinfo, gets an empty string back for the six that were hidden, and throws trying to convert it to an integer.

The tell is that it killed the CPU plugin as well as the GPU one. A GPU passthrough fault does not break CPU inference. I read past that for longer than I would like because I had already decided what the problem was.

What I’d tell someone

Check what you asked for before you upgrade what’s answering. The track list is one line and it was the entire fault. I was three clicks from buying hardware to fix a config filter.

A clean grep is not a clean system when the logs don’t go where you’re looking. Find out where the application actually writes before you trust the absence of errors.

When a component fails, check whether it failed in a way that component can. The CPU detector dying was proof the GPU was not the problem, and it was in front of me the whole time.

And the cat thresholds stay provisional for a week. If the misses cluster on small, occluded or night-time frames, that is what a better model fixes and the export work is justified. If daylight detection is fine, it is not. That decision gets made on a week of events, not on a forum thread — which is roughly how this whole thing started.