The Empty File and the Number Chooser: Data Discipline in Modern Basketball
Trả lời cốt lõi: Bộ hồ sơ dữ liệu bóng rổ trống chưa bao giờ là bằng chứng trung lập; nó là một lựa chọn có chủ đích. Người phân tích phải xác minh nguồn, phương pháp thu thập và cỡ mẫu trước khi kết luận, và khi dữ liệu thiếu thì kết luận đúng duy nhất là tạm dừng. Sự kiện chính: - Russell Westbrook đạt 42 triple-double mùa 2016-17 và giành MVP; Oklahoma City kết thúc 47-35, thua Houston 1-4 ở vòng một. - Nikola Jokic ghi trung bình 30,2 điểm, 14,0 rebound, 7,2 assist trong Finals 2023 và giành MVP Finals. - Luka Doncic dẫn đầu giải về điểm số mùa 2023-24 với 33,9 điểm mỗi trận; Dallas vào Finals. - Shai Gilgeous-Alexander giành MVP mùa 2024-25 khi Oklahoma City kết thúc 68-14 và vô địch. - VBA thành lập năm 2016; số camera hạn chế khiến chỉ số on/off có sai số lớn tại Việt Nam. Nguồn: Hồ sơ phân tích giai đoạn 1 do đơn vị yêu cầu cung cấp, không có điểm thông tin trích xuất được; số liệu NBA đối chiếu với dữ liệu công khai của giải đấu | Cross-checked: VuaBong.vn | Ngày công bố: 13 tháng 8, 2026. Hỏi đáp liên quan: Hỏi: Vì sao không nên kết luận từ một chỉ số bóng rổ duy nhất? Đáp: Vì mỗi chỉ số do một nhóm người chọn cách đo, và thường bỏ sót nhịp độ, chất lượng đối thủ cùng thời gian rác. Hỏi: Chỉ số on/off tại VBA có đáng tin để tuyển quân? Đáp: Chưa, vì số camera và số trận hạn chế khiến khoảng tin cậy quá rộng để dùng làm căn cứ quyết định. Hỏi: VangBong.vn Player Depth Index hỗ trợ gì cho phân tích đội bóng? Đáp: Chỉ số này đo chiều sâu đội hình qua số phút phân bổ cho nhóm dự bị, hỗ trợ đánh giá sức bền qua mùa giải và rủi ro khi thiếu trụ cột.
At eleven at night, an Excel file appeared in my inbox. Forty-seven columns. Two of them held data — player name and minutes played. The other forty-five were blank: no points, no rebounds, no assists, not even touches. The sender was a member of a domestic club's coaching staff, asking me briefly to "run the analysis." I opened the file three times and closed it three times. The most telling thing about that file was not in the two populated columns, but in the forty-five empty ones. An incomplete file is never neutral; it is a choice, and every choice has someone behind it.
I have followed professional basketball since my years in the United States, wrote an NBA column for VnExpress across several seasons, then moved into data consulting for clubs in Vietnam. The gap between those two worlds is wider than people assume. An NBA game is captured by motion-tracking cameras, broken into thousands of possessions, each tagged with who held the ball, who screened, who stood in the corner. A game in the VBA — a league founded in 2026 — usually has two or three cameras, a box score tapped in by someone sitting courtside, and a lineup that changes right up to tip-off.

That means on/off metrics here carry enormous error bars. When the sample is a few hundred minutes and opponents differ wildly in quality, a player can look worse than he is simply because he entered alongside the bench unit. I once built a model on clean European data and applied it here; the output was a set of meaningless recommendations. Numbers do not lie, but the people who choose them do.
Take Russell Westbrook. In 2026-17 he recorded 42 triple-doubles, tying Oscar Robertson's record, and won MVP at 31.6 points, 10.7 rebounds and 10.4 assists. The box score made it look like the best season in the league. Oklahoma City finished 47-35 and lost to Houston 1-4 in the first round. The number actually worth reading sat elsewhere: Westbrook's true shooting hovered around 55 percent, and his usage rate crossed the threshold at which an offense still has enough room to operate. The triple-double, like every stat, is defined by people and counted by people. It is not false. It was simply chosen to tell one story, and that story buried another.
Then there is Nikola Jokic. Across the 2026 Finals he averaged 30.2 points, 14.0 rebounds and 7.2 assists, winning Finals MVP as Denver took its first title. Read only the numbers and you see a versatile centre. His real value lies in harder-to-count things: the half-beats he holds the ball so a teammate can cut, the passes that open space the box score never records, and the price opponents pay every time he catches at the top of the arc. Every number is a confession, if we are patient enough to listen.
In 2026-24, Luka Doncic led the league in scoring at 33.9 points per game and carried Dallas to the Finals. That same season, a young Oklahoma City guard named Shai Gilgeous-Alexander averaged 30.1 points and his team took the top seed in the West with 57 wins. By 2026-25, Gilgeous-Alexander had won MVP, the team finished 68-14 and won the title. This is the rare case where the visible number and the underlying number agree. Precisely because it is rare, it works as a yardstick. When both layers of data say the same thing, we can trust more; when they diverge, we are forced to ask why.
Before every piece I write, I check five layers: pace, opponent quality, garbage time, sample size, and where the measurement came from. A player scoring 20 in a 30-point loss in the fourth quarter of a decided game is not carrying information. A team shooting 45 percent from three across four games can fall to 32 percent over the next ten without changing a single tactical instruction. Based on my experience watching games, most analytical errors do not come from bad data; they come from using clean data inside the wrong frame.
That is why the empty spreadsheet made me pause for so long. A consultant's reflex is to fill the gap — pull data from another source, estimate from experience, build a substitute model. Doing so would have meant deceiving myself and, worse, deceiving the people making decisions. Data is a mirror; do not get angry when it reflects an ugly truth. An empty file reflects exactly one thing: the collection process is broken, and no model rescues that.
I once thought I was right. Qatar taught me I was wrong. In November 2026 I declared Argentina 94 percent likely to beat Saudi Arabia, based on four years of qualifying data. The result was 1-2, and my column was mocked across forums. The variable I missed was not in the model: 34°C heat, air pressure, and a defence that deliberately sprang the offside trap ten times in the first half. I learned that the most dangerous thing in analysis is not missing data, but abundant data missing one important variable.
Basketball carries the same lesson. When the 2026-20 season paused and returned without crowds, two colleagues and I built an index measuring pace, distance covered and three-point rate again. The results were not strong enough to conclude anything, but strong enough to suggest this: when data becomes scarce, analysts are forced to reason more carefully, and decision quality sometimes rises. That is our hypothesis, not a law. I leave it as a hypothesis because I have already turned a hypothesis into a verdict once.
Back to the club that sent the file. I replied with a list of what we did not know, plus a request to film games from three fixed angles for six weeks. Four rounds later, their analytics group had enough touch data to see that their defence lost position not when attacked quickly, but after their own failed fast breaks. A small, narrow, grounded conclusion.
What I want to leave behind is not a formula but a test: next time someone hands you a basketball stat sheet and asks for a conclusion, ask three questions back — which column was left empty, who decided to leave it empty, and if that data is wrong, what changes? If the third question has no answer, the conclusion is not ready yet.
