Event Detection Is Easy. Editorial Relevance Is the Hard Part
This is one of the biggest misunderstandings in sports AI. People think once you can detect an event, you are close to solving highlights.

This is one of the biggest misunderstandings in sports AI. People think once you can detect an event,
you are close to solving highlights. You are not. Detection is the easier layer. Relevance is where
things get difficult.
A model can learn to recognise patterns associated with a shot, a tackle, a celebration, a set piece, or
a replay sequence. Fine. But editorial value sits above detection. A routine event can be irrelevant. A
scrappy near miss can be hugely important. A crowd reaction or a player’s body language can
matter more than the basic label attached to the action.
That is why purely event-based systems often feel disappointing in practice. They produce volume,
not judgement. They know something happened, but they do not always know whether the moment
deserves attention right now.
This is where richer video understanding starts to matter. Research such as InternVideo2 is
interesting because it moves closer to context rather than just classification. But even then, there is
no shortcut to editorial logic. The model can assist. It still does not become the producer.
A realistic workflow accepts that. Use event detection as the first filter. Use tracking, pose, and clip-
level understanding to improve ranking. Then keep a human in the loop for final relevance. That is
not a compromise. That is the right division of labour.

