Taekwondo CV
Real-time strike detection and scoring across four ringside cameras
A computer vision system that watches a taekwondo bout on four cameras at once, classifies what each fighter just did, turns it into points and shows the referee a live score. I built the ML side from scratch: the labelling schema and dataset, the detection cascade, the two-level deduplication that decides what counts as one strike, and the inference stack that runs it on a GPU in real time. Below is what was tried, what was thrown away, which thresholds the pipeline actually runs on and why they are asymmetric.
- RoleML/CV solo, plus architecture and backend
- InputFour live RTSP feeds, resized to 1280x720 before inference
- CoreTwo-stage YOLO cascade, 20 strike classes
- OutputScored events, annotated RTSP stream, live score over WebSocket
The constraint that shaped everything
The system had to answer the referee's question, not the detector's: which fighter, which technique, which target level, how many points. And it had to answer during the bout, not afterwards, on four simultaneous camera feeds, on one machine.
Two properties of the domain drove every later decision. First, the cost of errors is asymmetric: a missed strike is an argument at the end of the round, an invented strike is a wrong score on the board. Second, one strike is one event, but a detector fires on every frame it likes, so a kick that was visible for thirty frames must not become thirty events.

The pipeline end to end
- Capture: an RTSP feed per camera, decoded in its own thread, resized to 1280x720 before inference.
- Stage 1: a detector over the full frame finds the fighters and returns their boxes.
- Crop: each box is expanded by 20 px on every side and clipped to the frame.
- Stage 2: a 20-class classifier runs on that crop at 640 px and returns the technique, the target level and the fighter colour.
- Mapping: the class maps to an event code, and the points come from a table cached out of the database, not from the model.
- Deduplication: a per-class cooldown on the inference side, then a peak window on the event side.
- Persistence: the event, its confidence and a screenshot in object storage; the annotated frame goes back out over RTSP and the score over WebSocket.
Branch one: a single segmentation model
The first working version was one segmentation model over the full frame, with high-quality masks and a renderer that smoothed them over time: exponential blending with the previous mask, feathered edges, and a fade-out decay when a detection disappeared so the overlay did not flicker.
- The overlay looked genuinely good, which matters when it is going on a broadcast stream.
- Masks follow the body, so a limb crossing the opponent stays readable.
- One model, one set of thresholds, nothing to coordinate.
- Masks cost time per frame and bought nothing for the actual decision, which is a class, not a silhouette.
- At the far end of the ring a hand strike and a foot strike converge visually once the fighter is a couple of hundred pixels tall, and the single model had to localise and classify in the same pass. Classification is what degraded.
- The smoothing that made the overlay pleasant also delayed the moment a detection was considered gone.
Decision Kept in the codebase as a legacy single-model path behind a flag, because it is still the fastest way to debug a new dataset visually. Not the production route.
Branch two: pose estimation
The hypothesis was that joint geometry separates techniques better than appearance does: a head kick and a body kick differ in where the foot ends up relative to the hips and shoulders. I trained a pose model with a custom 35-keypoint skeleton and two classes, red and blue, on my own dataset, and exported it to OpenVINO.
- Keypoints are compact and interpretable, and a rule on top of them is easy to explain to a coach.
- Pose is far less sensitive to kit colour and hall lighting than raw appearance.
- There are too many legitimate ways to throw the same strike. The variance inside a class swallowed the difference between classes.
- In a clinch the keypoints of two fighters interleave and the skeleton fits the wrong body.
- It adds a model to the chain whose failures are harder to see than a bad box.
Decision Dropped. It is in this write-up on purpose: the architecture that ships is usually the one left standing after a few of these.
Branch three: the cascade that shipped
Splitting localisation from classification fixed the far-end confusion. Stage one only has to find people, which is the easy half and survives scale, and stage two sees a crop where the fighter fills the frame, which is exactly the condition under which a strike class is separable.
- Each stage gets its own threshold, and they can be tuned in opposite directions - see below.
- The hard problem shrinks: classification runs on a normalised crop instead of on a 1280 px frame.
- Stage one's boxes are useful on their own, so the overlay always shows the fighters even when no strike is being called.
- Cost is no longer constant per frame: stage two runs once per detected fighter, so a crowded frame is a slower frame.
- A stage one miss is unrecoverable - stage two never sees what was never cropped.
- Two models means two sets of weights, two exports and two things to version.
Decision This is the production path.

Thresholds, and why they point in opposite directions
The two stages are tuned against each other on purpose. Stage one runs permissive because a fighter it fails to find costs an event outright and there is no second chance. Stage two runs strict because an event it invents lands on the scoreboard, which is the expensive error in this domain. In effect stage one buys recall and stage two spends it on precision.
These are the defaults the service ships with; each hall and camera position gets them re-checked on recorded footage before an event, which is why every one of them is an environment variable rather than a constant.
- STAGE1_CONF
- 0.20
- Permissive on purpose: a fighter not found is an event lost.
- STAGE2_CONF
- 0.68
- Strict: a false class becomes a wrong score, so weak evidence is refused.
- IOU_THRESHOLD
- 0.45
- NMS overlap, shared by both stages.
- BBOX_PADDING_PX
- 20
- Crop margin, so a limb leaving the box is still inside the crop.
- stage 2 imgsz
- 640
- The crop is already tight; a larger input buys nothing but latency.
- RESIZE
- 1280x720
- Applied before inference - the single biggest throughput lever.
- cooldown
- 0.7 s
- Frame-level repeat suppression, keyed per class.
- peak window
- 2.0 s
- Event-level window, keyed per fighter, camera and round.
- A low stage-one threshold multiplies stage-two calls, so latency grows with the number of boxes rather than staying flat per frame.
- A strict stage-two threshold silently drops genuine but partially occluded strikes; that is a deliberate trade, not a free win.

Deduplication in two levels
The first level sits in inference: a cooldown per strike class, so the same technique from the same fighter cannot fire twice inside the window. The class already encodes the colour, so red and blue never suppress each other.
The second level sits in the event layer, and it is not suppression but arbitration. Within a window per fighter, camera and round, the system keeps the most valuable strike and removes the weaker one it had already written: a mid kick followed by a head kick followed by another mid kick resolves to the head kick. That matters because the model often sees the build-up and the peak of the same action as two detections, and the referee should see the peak.
- cooldown key
- strike class
- The class carries the fighter colour, so both fighters score independently.
- peak window key
- round + fighter + camera
- Arbitration is per fighter, not per frame.
- peak window rule
- highest points wins
- A better strike replaces the event already stored.
- A genuinely fast second strike inside the window is absorbed by the first. The window length is the knob that trades duplicates against missed follow-ups, and it is config, not a constant.
- Arbitration is per camera, so the same strike seen by two cameras still produces two events. Cross-camera fusion is the obvious next piece of work.

A bug worth keeping in the write-up
Event screenshots had their own cooldown, keyed on the camera and the event code. When red and blue landed the same kind of strike within a second, only one screenshot was taken and the second event was stored with an empty image URL. Events are created per fighter; the screenshot key was not. Adding the fighter to the key fixed it.
The general lesson is worth more than the fix: every deduplication key has to match the granularity of the thing it is protecting, and the easiest way to get that wrong is to copy a key from a neighbouring code path.
Cheap accuracy: the ring polygon
Not every person in the hall is in the bout. Spectators, coaches and athletes warming up in the background all generate perfectly valid detections. The system loads a ring polygon per camera and ignores anything outside it, and I wrote a small interactive tool that builds that polygon by clicking the corners on a live frame and saves it as JSON.
- Removed a large class of false positives for almost no compute.
- It is a rule, not a model: explainable, instantly adjustable, no retraining.
- Manual setup per camera, and it silently goes wrong if someone nudges the camera between sessions.
- A fighter at the very edge of the mat can fall outside a tight polygon.
How speed was measured
Latency claims are worthless without a method, so I wrote a harness that runs the candidate models over the same recorded bout and reports per-stage timings rather than one number. It discards warm-up frames, synchronises the GPU around each timer so it measures the work and not the queue, and prints average, median, p95 and p99 per stage plus the pipeline FPS and the ratio against the source frame rate - the answer to the only question that matters, which is whether it keeps up.
- Same video, same thresholds, one model swapped at a time.
- Warm-up frames excluded; the GPU is synchronised around every measurement.
- Per-stage avg, p50, p95 and p99 - tail latency is what drops frames, not the mean.
- Pipeline FPS against source FPS, printed as a real-time factor.
- The annotated video is written out too, so a speed regression and a quality regression are visible in the same artefact.
Decision The absolute numbers belong to the GPU they were measured on, so what I carry between projects is the harness, not the figures.
Inference and operations
- TensorRT engines are built from the weights on first run and cached per GPU model, so a machine change rebuilds instead of silently running a mismatched engine.
- Half precision by default, 1280 px engine input, with an automatic fallback to plain weights when the GPU or TensorRT is unavailable.
- CPU or GPU is chosen by one environment variable, so the same compose file runs on a laptop and on the venue server.
- Four camera workers in parallel, each with its own capture and inference loop.
- Annotated video republished over RTSP; score and match control over WebSocket.
- Event screenshots to object storage, metrics to Prometheus and Grafana.
Data, labelling and the split
The dataset is frames sampled across many bouts rather than consecutive frames of a few. The labelling schema carries attributes beyond the class: the fighter colour, whether the frame is the contact frame, the camera, and a quality flag with three values - clear, partial, doubtful.
That quality flag is the part I would build first again. It lets ambiguous frames stay in the dataset as data about ambiguity instead of either poisoning the training set or being silently deleted.
The split is by bout, never by frame. Frames from one bout share lighting, mats, fighters and camera angle, so a frame-level split leaks the test set into training and returns metrics that look excellent and predict nothing.
The referee is inside the loop
Every automatic event is stored with its confidence. The referee can add an event the system missed and remove one it invented, and those two actions are kept as explicit flags with a soft delete rather than as silent overwrites.
So the same table is both the match record and a labelled history of the model's false negatives and false positives, with the confidence that produced each one. That is the piece that transfers to any human-in-the-loop product: corrections only become training data if they are captured as structured data at the moment they are made.
Honest limitations
- No person tracking yet. Fighter identity comes from the class, and the fallback heuristic splits by frame half, which is fragile the moment the fighters cross over. Tracking is the next thing I would add.
- No cross-camera fusion: four cameras that all see one strike produce four events, and only the peak window per camera holds it down.
- Thresholds are re-checked by hand per hall; there is no automatic calibration against a small labelled sample.
- The dataset is small by CV standards, so per-class metrics on rare techniques carry wide error bars and I report them as such.

What I take from it
The interesting engineering was not the model. It was the definition of an event, the deduplication keys, the asymmetric thresholds, the split by bout, the measurement harness and the decision to store every correction. The detector is replaceable - that scaffolding is what makes the next detector measurably better than the last one.
Video
A clip of the system running is being added here.
Stack
- Python
- PyTorch
- YOLO / Ultralytics
- OpenCV
- TensorRT
- OpenVINO
- FastAPI
- PostgreSQL
- Redis
- Celery
- WebSocket
- MediaMTX / RTSP
- MinIO
- Docker
- Prometheus
- Grafana