Skip to content

InpaintNet fills gaps longer than its window with a fixed pattern. Also, output gives no way to easily filter it. #22

Description

@ahalp90

Hi, thanks for TrackNetV3. I've been using it with the released checkpoints to track
badminton shuttles in long broadcast videos. It's great.

But I've been using it so much that I've found a few bugs/possible improvements along
the way. This one was really cryptic, so I thought it might be helpful to share.

Please excuse the AI-generated text below. I've spent my morning exploring this bug and
have manually verified it, but Claude did most of the legwork so it seemed suitable to
use its write-up too.

Note: I ran this with --eval_mode nonoverlap, where frame windows tile (ie. stride=8).
I'm currently running a pass with stride=1 sliding windows and will update once I've got
the results. But given it's in the weights themselves I expect there'll still be some
signature remaining. Just the overlapping windows might smear it into a near-constant
position rather than the loop I've currently got.

The short version. When tracking loses the shuttle for a whole InpaintNet window, the
*model gets no evidence at all. Its input is all-zero coordinates plus an all-ones mask.
*Its output is then fixed by the trained weights alone, so every such window, in every
*video, gets the same 16 invented positions. predict() marks them Visibility = 1.

Because the inference path never outputs the Inpaint_Mask, downstream code cannot
easily separate them from real detections.

Why it happens. The InpaintNet interpolator's training masks are random per-frame
*dropout. Each frame is masked independently with probability 0.3 (get_random_mask,
*np.random.binomial(1, mask_ratio, ...) in train.py), and the loss is computed only at
*masked positions. Under that recipe a fully masked 16-frame window almost never occurs
*(0.3^16 probability, about 4 in a billion), so the model never learns a "no evidence"
*behaviour.

At inference, generate_inpaint_mask marks whole gaps for filling with no length cap;
it only checks the y-coordinate of the two frames flanking the gap. A long gap therefore
hands the model exactly the input it never saw in training, and the output is whatever
the weights default to.

You can see it without a video:

ckpt = torch.load('ckpts/InpaintNet_best.pt', map_location='cpu') seq_len =
ckpt['param_dict']['seq_len']

net = InpaintNet() net.load_state_dict(ckpt['model']) net.eval()

coords = torch.zeros(1, seq_len, 2)   # a window inside a long gap
mask = torch.ones(1, seq_len, 1)
with torch.no_grad():
    out = net(coords, mask)[0]

print('x px:', [int(v * 512) for v in out[:, 0]])
print('y rows:', [int(v * 288) for v in out[:, 1]])

With the released checkpoint this prints, give or take a pixel of CPU-vs-GPU rounding:

x px:   [240, 244, 244, 245, 242, 245, 244, 244, 245, 246, 242, 240, 241, 241, 241, 247]
y rows: [81, 73, 70, 69, 67, 68, 69, 70, 68, 68, 71, 73, 77, 80, 83, 84]

All our saved tracks have that little bobbing loop wherever a whole window had no
detection. In our three broadcast videos it covers about a third of all frames, much of
it inter-rally footage where there is genuinely no shuttle to track.

Three suggestions. I think any of these would help:

  1. Save out the inpaint mask at inference, as a CSV column or a supporting file, so
    users can choose how they use it. As shipped it never reaches the output: the only
    save_inpaint_mask path is the ground-truth-dependent training-data writer.
  2. Skip filling with an explicit sentinel value where there is nothing to interpolate from:
    for example, leave a gap unfilled once it spans a full window, or skip windows
    containing no real detection. Interpolation is not defined there anyway.
  3. If retraining is ever on the table, the realistic masks already exist in the pipeline:
    generate_mask_data.py builds gap-shaped masks from real TrackNet misses, and the
    dataset loads them, but train_inpaintnet binds them to _ and draws the random
    masks instead (they are only used at validation, for checkpoint selection). Mixing
    the real masks into training would show the model contiguous gaps of the shape
    inference actually asks about.

I'm happy to open a PR for the code to save the Inpaint mask table.

Thanks again for the project.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions