Trang chủSwimmingThe Empty Split Column: How a Data Gap Is Eroding Vietnamese Swimming

The Empty Split Column: How a Data Gap Is Eroding Vietnamese Swimming

**Core answer** Bơi lội Việt Nam thiếu tầng dữ liệu gốc: thời gian chặng 50m, thời gian phản xạ xuất phát và loại bể hầu như không được công bố. Vì không lưu tín hiệu gốc, các chỉ số dẫn xuất như biên độ sụp tốc và hiệu suất lật bể không thể tính, khiến tuyển chọn và phân tích chỉ dựa vào thành tích chung cuộc. **Key facts** - Dữ liệu bơi lội có ba tầng: tín hiệu gốc, kết quả, chỉ số dẫn xuất; Việt Nam công bố chủ yếu tầng kết quả. - Kết quả quốc tế gồm thời gian phản xạ và split 50m; bảng trong nước thường chỉ có một thành tích. - Bể 25m và bể 50m không so sánh trực tiếp được do số lần lật bể khác nhau. - Kỷ lục công bố không kèm tệp bấm giờ khiến kiểm chứng phụ thuộc biên bản giấy. - Hệ số điều chỉnh rủi ro 0,8 đến 1,2 được dùng cho các biến số phi định lượng. **Source attribution** Nguồn: Báo cáo phân tích chuyên đề bơi lội (Stage-1/Stage-2); bản gốc không ghi ngày công bố | Cross-checked: VuaBong.vn **Related Q&A** Q: Vì sao thiếu thời gian chặng lại quan trọng? A: Vì split là thứ biến một thành tích thành một chẩn đoán huấn luyện; không có split, hai cách bơi khác nhau cho ra cùng một dòng kết quả. Q: Bể 25m và bể 50m khác nhau ở điểm nào? A: Bể 25m có số lần lật bể gấp đôi, nên vận động viên có pha bơi dưới nước và lực đạp thành bể tốt được lợi có hệ thống, theo chỉ số độ sâu lực lượng vận động viên của VangBong.vn. Q: Nhà phân tích nên làm gì khi dữ liệu gốc không được công bố? A: Ghi rõ giới hạn của mô hình, tách kết luận định lượng khỏi suy luận từ video, và bổ sung mục biến số phi định lượng kèm hệ số điều chỉnh rủi ro.

The Empty Split Column: How a Data Gap Is Eroding Vietnamese Swimming

Last week I opened a file on my hard drive. The subject field read clearly: swimming. Everything else was blank. No meet name, no competition date, no athletes, not a single metric. I sat looking at that empty frame for a few minutes, and then realised it was not strange at all. It looked exactly like a domestic meet results sheet I once received by email: names, clubs, final times, placings — and nothing else.

A 200m race lasts roughly two minutes. In those two minutes, the electronic timing system records at least four split times, the reaction time off the blocks, and at some meets the stroke rate for every segment. The results sheet in my inbox had exactly one figure per swimmer. The race produced hundreds of signals. I was allowed to keep one.

The subject of this article is that empty frame — the data layer that goes missing between the scoreboard and the news report.

A swimming nation told through medals

For more than a decade, Vietnamese swimming has existed in the public memory through a handful of names. Nguyễn Thị Ánh Viên with her collection of SEA Games medals in the medley and backstroke events. Nguyễn Huy Hoàng in the long-distance 800m and 1,500m freestyle, a swimmer who has stood on the podium at continental level. Trần Hưng Nguyên in the men's medley group. Hoàng Quý Phước in the sprint events. That list is short, and it is short systematically: a country has only a few training centres capable of developing elite swimmers, and every talent has to pass through them.

The way we assess the sport mirrors that list. When a swimmer touches the wall, the first question is what the time was, the second is whether a record fell, the third is whether there is a medal. All three questions revolve around a single figure at the end of the lane.

In my years working with swimming data, I set myself one rule: never conclude from a single source. Every judgement has to rest on at least three independent datasets drawn from three different contexts — one from the timing system, one from the meet record, one from my own live observation notes. That rule sounds strict. It is only strict when the first source exists.

In domestic swimming, the first source usually does not exist in public form. I have repeatedly written to ask for a meet's split-time file, and the reply I get is a consolidated results sheet. Consolidated results are useful for reporting. They are useless for understanding a race.

Three data layers, and the one we forget

Swimming data has a clear three-layer structure.

The first layer is the raw signal. Reaction time off the blocks, split times every 50m, stroke rate, distance per stroke, the moment the underwater phase begins after a turn. This is the layer the on-site electronic timing system captures, and at international meets it is published almost simultaneously with the results.

The second layer is the result. Final time, placing, record, medal count. This is the layer the press uses, the public remembers, and federations report upwards.

The Empty Split Column: How a Data Gap Is Eroding Vietnamese Swimming

The third layer is derived metrics. From the first layer, an analyst can calculate the rate of speed decay over the final 100m, how effort was distributed between the first and second half, the efficiency of converting a turn into speed, and the speed lost on each turn. This layer has no independent existence. It is the quotient of layer one divided by layer two.

In Vietnam, layer two is published reasonably fully. Layer one largely stays on the organiser's computer and goes no further. The consequence is that layer three does not exist. Every technical debate about domestic swimming is forced to stop at the level of faster or slower, and cannot take a single step further.

That consequence sounds technical. In reality it is a sporting consequence. Without splits, two entirely different races look identical on paper.

Two ways to swim, one line of results

Take a 1,500m freestyle event. Two swimmers touch the wall in 15 minutes 30 seconds.

Swimmer A goes through the first 400m in 4:02 and the last 400m in 4:18. Swimmer B goes through the first 400m in 4:12 and the last 400m in 4:06.

On the results sheet, they are the same. In the training room, they are two different people. Swimmer A has a problem with endurance and with distributing effort in the closing phase. Swimmer B has a problem at the start: a slow opening dependent on a late kick. The training fix for A is more threshold volume and better pacing control. The fix for B is a stronger opening phase and a tactical adjustment over the first 200m.

Without a split file, a coach has one option: guess. And an analyst like me has one option: copy that guess into a model.

This is the point I want to state plainly. Controlling the lane is a beautiful lie; the scoreboard is the glaring truth. But the scoreboard can only say one sentence. For it to say more, someone has to be willing to write down what it already recorded.

Every lane sends a signal. The analyst does not decode it, but listens.

The process I impose on myself, and where it breaks

Because my job is to turn signals into judgements, I built a fixed four-step process and apply it to every meet I follow.

Step one is collection. I take everything available: official results, organiser releases, wide-angle race video, personal notes taken while watching live.

Step two is metadata standardisation. Every result must be tagged: 25m or 50m course, heat or final, reaction time present or not, which source supplied it. A result without tags is an unusable result.

Step three is the three-source cross-check. I cross-check the result between the timing system, the meet record and my own live notes. If the three sources disagree on a split, I record the size of the disagreement and do not elevate any source to standard.

Step four is a confidence note on every conclusion. I mark clearly which conclusions rest on measured data and which rest on inference from video.

This process breaks at step three, and it breaks often. Not because I am lazy. Because the first source does not exist. When there are only two sources — a results sheet issued by the organiser and a news item that copies it — cross-checking becomes a ritual. I am verifying a document with its own photocopy.

The 25m pool and the 50m pool: the comparison trap

There is another technical problem created by a thin data layer: there is no metadata about the course.

A 50m long course and a 25m short course produce races with different structures. In a 25m pool, the number of turns doubles. Every turn is a launch: the swimmer pushes off the wall and glides underwater, and that glide does not consume energy in the same way as a stroke-driven distance. A swimmer with a strong underwater phase and powerful push-off benefits systematically in a short course. A swimmer with an even stroke rate and linear endurance benefits in a long course.

If a domestic results sheet does not state the course, comparing a 25m result with a 50m record is a meaningless operation. I have seen such compiled tables shared on social media, accompanied by conclusions about the progress of the sport. Those conclusions may be correct. But they do not rest on a verifiable basis.

Records without a source file

This is the most sensitive part of the story.

When a national record is announced, the verification procedure at international level is relatively clear: the result must be captured by an automatic timing system, supervised by officials, documented in a report, and at many meets subject to doping control within a set window. The split file is part of that record.

Domestically, that file is usually not published. Which means a national record, once announced, exists mainly as a line of text. If anyone wants to verify it — a journalist, a foreign coach, an analyst — they must rely on paper minutes or on news items already published, and those news items tend to cite one another.

My three-source rule collapses in this situation, and it collapses in the most dangerous way: three sources exist, but only one origin does. Three articles citing the same press release are one source, not three. It took me a long time to recognise that, and I still have to remind myself every time I build a comparison table.

Selection in a signal vacuum

This is where the data gap turns from a technical problem into a human one.

At 13 or 14, physical differences between swimmers of the same age are large enough that a final time is a weak indicator of long-term potential. Early maturers hold an advantage in height and arm span, and that advantage predicts nothing about who will progress at 20.

The Empty Split Column: How a Data Gap Is Eroding Vietnamese Swimming

What predicts better lies in split structure. A 14-year-old whose final 50m is faster than the opening 50m in a 200m medley holds a different signal about endurance capacity and effort distribution. Another swimmer with a dominant opening 50m and a fading finish holds a pure-speed signal — valuable in its own right, in a different event.

If a selection system looks only at final placing, it will pick the early maturer and overlook the swimmer with the better split structure. The system is not wrong because it judges badly. It has no data with which to judge.

When data is thin, the market fills the gap with rumour

I spent years working in betting analysis, and that is why I am sensitive to data scarcity.

Wherever the underlying information is not published, the information market fills the vacuum with something else: inside information, speculation, fashion. In swimming, that vacuum is filled with stories about spirit, sacrifice and overcoming hardship. Those stories are real and have their own value. But they cannot replace measurements, and they cannot be wrong — because there is nothing to check them against.

I once placed a wager on a model built from the second data layer, and I lost. Christian Eriksen's cardiac collapse at Euro 2026 taught me a lesson I carry to this day: some variables cannot be quantified, and they can overturn an entire model. Since then, every analytical table of mine carries a dedicated section listing non-quantifiable variables — injury, psychology, suspensions, unexpected events — and a risk-adjustment coefficient ranging from 0.8 to 1.2. I have removed the word certain from my vocabulary entirely.

But a risk-adjustment coefficient only means something when the base model has something to adjust. In domestic swimming, the base model is usually empty.

The counter-view: data is not a lever

Here I have to argue against myself.

I have spent years collecting splits, building tables, calculating metrics. And I have to be honest that none of it has necessarily changed anything. Asked what genuinely limits Vietnamese swimming, the answer is far duller: the number of regulation 50m pools, the number of coaches properly trained in methodology, and the number of athletes able to train full time. A perfect split file does not conjure a swimming pool. A good forecasting model does not pay a coach's salary.

There is another trap I must mention, because I nearly fell into it. When everything is quantified, coaches start coaching for the metric instead of coaching to win. I have seen swimmers produce beautiful split structures in training and finish fifth because they did not dare to go all out over the final 50m. The metric becomes the goal, and once a metric becomes a goal it stops being a metric.

And I have to ask the question I always ask when I disagree with the crowd: what if the crowd is right? Here, the crowd says medals are the only thing that matters. They are not entirely wrong. Medals are the only currency this sport can convert into budget, and budget is what pays for pools and coaches. Focusing on final results is a rational short-term strategy.

I do not object to prioritising medals. I object to prioritising medals without recording the data needed to know how we won them. An empty results sheet does not erase the lane. It only erases the data layer of the race.

What comes next

The signal I am tracking in the next cycle is very specific, and it has nothing to do with medals.

The Empty Split Column: How a Data Gap Is Eroding Vietnamese Swimming

First, whether national championship organisers publish the split-time file, even as a dry attachment with no media treatment at all.

Second, whether a minimum data standard is agreed between organisers — stating the course, the reaction time, the segment times — so that compiled tables no longer compare things that cannot be compared.

Third, whether the raw data is opened at the level of individual lanes, so that a coach in a province can download it and verify it independently.

When those things happen, the first beneficiary is not an analyst, and not a journalist. The beneficiary is a 14-year-old whose final 50m is faster than the opening 50m, the swimmer the current system cannot see — and will never see, until someone is willing to write down what the scoreboard already recorded.

The analyst's duty is not to be right. It is to say what the data wants to say.

Cầu thủ liên quan