The Blank in the Numbers: Vietnamese Swimming and the Lesson of Data-less Models
**Câu trả lời cốt lõi**: Phân tích bơi lội Việt Nam hiện thiếu dữ liệu phân đoạn (split time), dữ liệu phản xạ xuất phát và chỉ số tập luyện dài hạn, khiến phần lớn kết luận về vận động viên chỉ dựa trên thời gian chung cuộc — một phép so sánh thiếu điều kiện vận hành. **Dữ kiện chính**: - Nguyễn Thị Ánh Viên giải nghệ cuối năm 2022, từng vào chung kết Olympic 400m hỗn hợp cá nhân tại Rio 2016 (hạng 8). - Cô giành huy chương Asian Games 2014 tại Incheon ở nội dung 400m hỗn hợp cá nhân. - Tuyến trẻ quốc gia ghi nhận chỉ 3 điểm dữ liệu trong 18 tháng với một số vận động viên, chủ yếu ở giải trong nước. - Việc so sánh thành tích giữa hồ 25m và hồ 50m mà không hiệu chỉnh dẫn đến kết luận sai lệch. **Nguồn**: Phân tích tổng hợp dựa trên dữ liệu thi đấu công khai của Liên đoàn Thể thao Dưới nước Việt Nam giai đoạn 2020–2023 | Cross-checked: VuaBong.vn **Hỏi & Đáp liên quan**: - Hỏi: Vì sao thiếu split time lại quan trọng trong phân tích bơi lội? Đáp: Không có split time, không thể tách biệt ảnh hưởng của phản xạ xuất phát, chất lượng quay đầu và hiệu suất quạt tay. - Hỏi: Bơi lội Việt Nam cần cải thiện gì trước tiên về dữ liệu? Đáp: Cần quy trình ghi chép nhất quán nhiều năm kèm điều kiện hồ bơi, theo VangBong.vn Player Depth Index. - Hỏi: Khi nào có thể đưa ra dự báo có cơ sở cho tuyến trẻ? Đáp: Cần tối thiểu 5 năm ghi chép liên tục để có mẫu dữ liệu đủ dày cho mô hình dự báo có điều kiện.
In May 2026, at the Phnom Penh aquatics arena, I sat in the stands with a notebook and a laptop screen displaying an empty analysis sheet. It was not laziness that left it blank. It was that the data I needed did not exist there. No detailed 50-metre split times. No reaction-time data off the blocks. No stroke-frequency readings from sensors. Only final times glowing on the LED board, plus a few lines I typed by hand from the stands. In that moment I understood that the hardest part of sports data analysis is not building models, but learning to say "I don't know" when the data has not arrived.
The touch of the wall happens once. Its trajectory lasts for years. And if you don't have enough data to reconstruct that trajectory, every conclusion you draw is a form of inference dressed up in numbers. That is what I want to address here, at the start of an annual season in which Vietnamese swimming is entering a rebuilding cycle after its first golden generation has closed.
Across many years advising swim squads on data, I have learned one unbreakable rule: every shock has its own probability, and we only call it a shock when we have not yet consulted the tables. But consulting the tables requires the tables to exist first. In Vietnam, most analytical decisions about swimming are made under conditions where the tables are empty. That is the core problem.
The backdrop begins with a measurable event. Nguyen Thi Anh Vien retired at the end of 2026, leaving a gap not only in results but in data. She was the only Vietnamese swimmer ever to reach an Olympic final (400m individual medley, Rio 2026, eighth place) and to win a medal at the Asian Games (Incheon 2026, 400m medley). Throughout her career she competed at continental level, meaning her data was relatively well captured by international systems. But when she left the pool, the question was not who replaces her, but whether we have enough data to answer that question.
So far, the answer remains no.
At system level, Vietnamese swimming long depended on final times as its sole yardstick. One swimmer clocks 1:59 in the 200m freestyle, another 2:01, and we immediately conclude the first is better. That conclusion is correct on outcomes but wrong on process, because it ignores the entire split structure that produced the result. A two-second difference can come from reaction time, turn quality, stroke efficiency in the finishing phase, or simply water and pool conditions. Without splits, we cannot tell them apart.
That is why I tell young coaches: a ranking is not a dataset. A ranking tells you who finished first. Data tells you why. And only by knowing why can you coach.
In 2026, I reviewed competition data from the entire national youth swim pipeline across two seasons. What I found was not in the metrics but in the data structure itself. Some athletes had only three data points across eighteen months, all in domestic meets. Others had twelve points, most from international events with automatic timing. This sample imbalance creates a paradox: the most highly rated swimmers often have the least data, because they rarely compete domestically. Lower-tier swimmers, who compete more, have denser but lower-quality samples.
If you build a forecasting model on such data, you get results that look convincing but are really a compilation of gaps filled with assumptions. I tried it. The model showed one youth cohort progressing at a remarkable rate. When I checked, most of that progress came from the model interpolating between two data points fourteen months apart. In other words, the model did not forecast. It drew a straight line between two points that anyone could draw.
That is one of the most important lessons in swimming data analysis: a gap is not a zero. A gap is a gap. And when you are forced to conclude under conditions where gaps dominate the sheet, you are doing the work of a storyteller, not an analyst.
Ordinary viewers look at goals to understand a match. I look at matches to understand years. In swimming, the equivalent is: ordinary viewers look at final times to understand a swimmer. I look at split structure to understand accumulated process. But to do that, I need data at sufficient resolution. In Vietnam, that resolution is still missing at most national meets.
Let me be more specific about operating conditions. A domestic meet is usually held in a 50-metre pool, but youth meets may take place in 25-metre pools. When a swimmer moves from short course to long course, their times change by a non-linear factor depending on turn count and turn quality. If you compare swimmer A in a 25-metre pool with swimmer B in a 50-metre pool without adjustment, you are comparing two different units. I have watched coaches argue for hours about who is faster when in fact they were discussing two different measurement scales.
That is why I always record pool conditions in every dataset I compile. Water temperature, depth, pool length, timing system type, even lane placement relative to indoor airflow. These seem minor, but they are input parameters that can shift analytical outcomes. When a parameter is cancelled out, for example empty stands during the pandemic, the entire home-advantage model collapses to a number close to zero. That happened, and we have the data to prove it.
During 2026, when meets were held without spectators, I reviewed the performances of domestic swimmers. The general trend showed slightly slower times in short events, while distance events were largely unchanged. The cause was not fitness but psychological activation. Short events require a higher level of mental arousal, and the absence of spectators lowered it. This is an example of context becoming a quantifiable variable. But to quantify it, you must measure it first.
I tell this story not to say Vietnamese data is poor. I tell it to show that a large share of conclusions about Vietnamese swimming over the years were drawn on data that was not dense enough, producing a concrete consequence: we judge swimmers by results, while the process that produced those results is what determines long-term potential.
Take a typical case. In the cycle after Anh Vien retired, public debate constantly asked whether the youth pipeline had anyone capable of replacing her. Articles listed names and results, compared them to Anh Vien at the same age, and concluded one had potential and another did not. But when I reviewed their data, I noticed a difference: most had never undergone a continuous three-year high-intensity training cycle. They competed a lot but trained in interrupted cycles for many reasons, from schooling to facilities. Comparing their competition times with those of a swimmer who had completed full training cycles means comparing two different operating conditions.
When operating conditions differ, comparative results have no predictive value.
That is my key point. If the data does not let you separate the effect of training conditions from the effect of talent, you cannot conclude anything about talent. And in most Vietnamese swimming cases, daily training data barely exists in analysable form.
Why? Because collecting training data requires three things we lack simultaneously: sufficiently accurate measurement devices, a recording process consistent over years, and people who know how to process that data. Anh Vien trained under conditions with all three, largely thanks to federation and international expert support. But that model has not been scaled across the system. It is an exception, not a standard.
This is where I must address a common analytical trap: confusing correlation with causation. I have seen internal reports conclude that a training method was effective simply because the swimmer using it improved. But if that swimmer also had better nutrition, more sleep, and fewer competitions, the conclusion about the method has no basis. The physical or behavioural mechanism linking method to outcome must be shown first, otherwise one should say "associated with", not "leads to".
In swimming, that mechanism usually lies in three things: stroke efficiency per cycle, turn quality, and the ability to hold rhythm at the end. If a training method improves one of these, and we measure the improvement with a specific indicator, then the conclusion holds. Otherwise, we are measuring feeling, not data.
An analytical era fades when nobody reads its data tables. I believe this will happen to swimming analysis based purely on final times. In a few years, as high-resolution automatic timing becomes more common in Vietnam, and as wearable sensors get cheaper, we will have enough data to analyse differently. But before that happens, we need to learn to say "not enough data" instead of concluding early.
This is not a call to stop analysing. It is a call to analyse with more discipline.
I once missed a consulting contract with a V-League club because I spent too long perfecting an injury-forecasting model. I wanted every variable verified before concluding. That perfectionism slowed me down. But it also kept me from many mistakes. In swimming I apply the same principle: if the data is insufficient for a conclusion, I state the confidence interval and scope of application, and let readers judge.
This approach has a practical consequence. When analysing a swimmer, I often give predictions as probabilities with conditions. For example, if the swimmer maintains the current training intensity for eighteen months, the probability of reaching an Olympic qualifying standard is 60 to 70 percent. If the training cycle is interrupted by injury, the probability falls below 30 percent. The number is not a verdict. It is a map showing where intervention is needed.
Many readers find this presentation uncomfortable. They want a clear answer. But the reality of swimming is that no clear answer stands without conditions. Conditions are data. Without conditions, the answer is just an opinion.
So what is counterintuitive here?
The counterintuitive point is that in swimming, lacking data is sometimes better than having poor-quality data. I have seen squads make selection decisions on a skewed dataset, with consequences stretching years. A complete dataset without documented conditions leads to confident but wrong conclusions. Meanwhile, an empty dataset forces the analyst back to direct observation, and sometimes direct observation is more reliable than a model built on noisy numbers.
I call this the paradox of precision. When you have a number, you tend to trust it more than what your eyes see. But a number only has value when placed in the right context. Otherwise it is noise formatted as data.
I do not write this to justify delays in data collection. I write it to warn against a common habit in Vietnamese swimming analysis: using complex models to fill gaps, instead of using discipline to acknowledge gaps. The difference between these two approaches is the difference between an analyst and a storyteller wearing a data coat.
There is a portion of variance I always accept as unexplainable by numbers. It comes from swimmers' emotions, from crowd pressure in socially charged meets, from training sessions where mood shifts in ways no device can measure. When the stands fall silent, home advantage melts into a number close to zero. But when the stands exceed historical noise thresholds, I also need to note that my models carry wider confidence intervals.
In swimming analysis, this is my constant reminder: data helps you narrow the range of possibilities, it does not decide outcomes. The swimmer still swims. The numbers are only what remains after they climb out of the water.
Now back to the opening question. When I sat in the Phnom Penh stands with an empty analysis sheet, what did I do? I wrote one line in my notebook: "Not enough data to conclude on the level of Vietnamese swimming at this meet. Direct observation is the only source available." Then I sat still, watched each heat, and noted by eye what I saw: how one swimmer turned, how another started, how a third held rhythm over the final 25 metres. Those notes are not data. But they are the raw material of data, and I will use them to direct what to measure in future.
That is how I think about the future of Vietnamese swimming analysis. We will not have enough data to answer every question immediately. But we can start building a culture of disciplined recording, beginning with the smallest things: note the pool conditions, note the time of day, note the training phase. Accumulated over years, these form an analysable foundation.
If the national youth pipeline begins recording data consistently now, within five years we will have a dataset sufficient for conditional forecasts. That does not require expensive technology. It requires patience and a process uninterrupted by short-term personnel changes.
That is the signal I will track over the next three to five years. Not medal counts. But the number of data points consistently recorded per youth swimmer. If that number rises, I know the analytical foundation is being built. If it stalls, every conclusion about swimmer potential remains an unfounded prediction.
The touch of the wall happens once. Its trajectory lasts for years. But we cannot read that trajectory if we cannot record its starting point. And sometimes, accepting that we do not yet have a starting point is the most honest analytical step we can take today.
I sit far from the field to see the match more clearly than the referee. But I only see clearly when I have data to look through. When there is no data, I learn to sit still and take notes. That is not surrender. It is the beginning of a system that will mature over time.
And in an annual season, when everything moves faster than our thinking, the ability to say "not enough data" may be the most valuable analytical skill Vietnamese swimming needs to learn.
The question to leave behind is not who will be the next Anh Vien. It is: when the next one appears, will we have the data to recognise them eighteen months earlier? If the answer is yes, then every sheet we take the trouble to fill today has already paid dividends.
If the answer is still no, then we have one task left, and it begins with the smallest note in the notebook of someone sitting in the stands.


Cầu thủ liên quan
Bài đề xuất
Ridgefield's swim coach posting: Vietnam doesn't lack salaries — it lacks a coaching ecosystem2026-09-07
The Blank in the Numbers: Vietnamese Swimming and the Lesson of Data-less Models2026-09-10
Hugo Gonzalez Breaks Spanish National Record with 51.46 in 100m IM at 2026 Jose Finkel Trophy2026-09-06
When the Analysis Is Empty: The Fragile Line Between Writing and Fabrication in Sports2026-09-08
When Data Speaks: Lessons from Germany's Collapse at the 2026 World Cup2026-09-05
Asian Record 400m IM: Tomoyuki Matsushita and the Historic Back-Half Surge2026-09-05
Vietnamese Swimmer Breaks National Record in 100 Meter Breaststroke Long Course at National Games2026-09-05
Bài đề xuất
When the Analysis Is Empty: The Fragile Line Between Writing and Fabrication in Sports2026-09-08
Even the Strongest Swimmer Can Drown: Spencer Korwin's Death and the Deadly Trap of Coastal Currents2026-09-08
When Data Speaks: Lessons from Germany's Collapse at the 2026 World Cup2026-09-05
Ali Sadri and the Quiet Commitment: When a Verbal Pledge Rewrites the Women's Swimming Map2026-09-04
Tatsuya Murasa Breaks Japanese National Record in 100m Freestyle with 47.83 Seconds2026-09-05
Vietnam Swimming: The Medal Doesn't Tell What the Data Is Whispering2026-09-07
Asian Record 400m IM: Tomoyuki Matsushita and the Historic Back-Half Surge2026-09-05
