Input resolution controls how much visual detail reaches an object detector, but it does not predict end-to-end latency by itself. In our six-configuration Apple M4 Pro test, higher resolution consistently improved COCO accuracy while p95 latency did not increase monotonically at every step.
What did we test?
We evaluated official YOLO11n and YOLO11s checkpoints at 320, 480 and 640-pixel square inputs on all 5,000 images in the COCO 2017 validation split. Each configuration then ran 20 warm-ups and 200 timed, batch-one end-to-end predictions on a deterministic image sample. We report p50, p95 and p99 rather than relying on an average. We also repeated the 640-pixel validation and timing on one NVIDIA L4 in Cloud Run.
How much accuracy did resolution add?
YOLO11n rose from 28.82% mAP50–95 at 320 pixels to 38.88% at 640 pixels, a 10.06-point gain. YOLO11s rose from 37.52% to 46.37%, an 8.85-point gain. Person-class AP showed the same direction, reaching 57.94% for YOLO11s at 640.
Did smaller inputs always run faster?
No. YOLO11s measured 10.09 ms p95 at 320, 10.20 ms at 480 and 13.55 ms at 640. YOLO11n measured 12.73, 13.71 and 10.32 ms respectively. Metal execution, preprocessing, postprocessing and system scheduling make the relationship non-linear, which is why each export and target device needs direct measurement.
Which configuration did the frozen rule select?
Before running the experiment, we defined the candidate as the lowest-p95 local configuration within 2.0 absolute mAP50–95 points of YOLO11s at 640. Only the 640-pixel YOLO11s point cleared that 44.37% floor, so it remained the selection at 46.37% mAP50–95 and 13.55 ms p95.
What changed on Cloud Run L4?
At 640 pixels, the L4 PyTorch/CUDA run measured 13.05 ms p95 for YOLO11n and 13.22 ms for YOLO11s across 200 batch-one observations each. The similar end-to-end tail latency, despite YOLO11s doing more inference work, shows why preprocessing, result construction and synchronization matter at batch one. Because the operating system, runtime and hardware differ from the M4 test, this is deployment evidence—not a controlled accelerator ranking.
What does 13.55 ms mean for camera capacity?
Using half the p95-derived throughput as headroom gives 36.9 aggregate frames per second. Mechanically, that is 18 cameras sampled at 2 fps, seven at 5 fps, or three at 10 fps. These are sizing hypotheses: camera decoding, queues, network transport, application I/O and reject-control timing still consume budget.
Why is this not proof of defect-detection performance?
COCO contains generic objects, not a plant's defects, optics, material variants or changeovers. The benchmark answers a hardware and configuration question. A production release still needs rights-cleared line images, defect-level recall, an explicit false-negative limit and a test across every operating condition.
Practical takeaway
Measure the exact model, input size, export and device; freeze the accuracy gate before seeing the timings; and size on tail latency with headroom. The smallest checkpoint or image is not automatically the best production choice.