envisionboxbio-apetrack demo: Zero-shot multi-ape tracking and pose estimation using OWLv2, ByteTrack and OpenApePose

Authors
Affiliations

Wim Pouw

Tilburg University

Nisarg Desai

Emory University

PyPI  ·  GitHub  ·  Zenodo DOI  ·  EnvisionBoxBio

Detection and pose estimation of apes in video without training or labeling (zero-shot).

Wim Pouw, Department of Computational Cognitive Science, Tilburg University

Nisarg Desai, Emory National Primate Research Center, Emory University

DOI

envisionboxbio-apetrack demo samples

Demo page: https://wimpouw.github.io/using_envisionboxbio-apetrack/

Pipeline: OWLv2 (boxes, prompted with ape names) -> ByteTrack (track IDs within a video) + stitching across short occlusions -> Savitzky-Golay smoothing bounding boxes -> OpenApePose (16 keypoints per ape) -> Savitzky-Golay smoothing smoothing keypoints -> CSV + labelled video.

The current tool serves as an easy to use “out of the box” pose tracker for apes. While there are many easy to use deep learning frameworks to train one’s own model with your own labeled data, to my surprise there were no easy to use ape-specific tracking that work on any video. The available software, like OpenApePose (Desai et al., 2023) for keypoint detection, were however there to be easily combined with powerful top-down bounding box detection models (Owlv2), which together with software for stable detection over sequences (ByteTrack), makes a promising pose estimation tool.

OpenApePose (Desai, 2023) was trained on 71,868 annotated images:

Please reach out to w.pouw@tilburguniversity.edu if you want to help out with proper evaluation against state of the art (but usually less user-friendly) computer vision models.

Citation

Citing envisionboxbio-apetrack

  • Pouw, W., Desai, N. (2026). envisionboxbio-apetrack: Zero-shot multi-ape tracking and pose estimation using OWLv2, ByteTrack and OpenApePose (Version 0.1.1) [Computer software]. Zenodo. https://doi.org/10.5281/zenodo.22992133

All versions: https://doi.org/10.5281/zenodo.22992132

Citing the components

This python package is a productive recombination of available and openly licensed software. Therefore first and foremost these packages need to be cited:

  • OWLv2: Minderer, M., Gritsenko, A., & Houlsby, N. (2023). Scaling open-vocabulary object detection. Advances in Neural Information Processing Systems, 36, 72983–73007. https://doi.org/10.52202/075280-3191
  • ByteTrack: Zhang, Y., Sun, P., Jiang, Y., Yu, D., Weng, F., Yuan, Z., … & Wang, X. (2022). ByteTrack: Multi-object tracking by associating every detection box. In European Conference on Computer Vision (pp. 1–21). Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-20047-2_1
  • OpenApePose: Desai, N., Bala, P., Richardson, R., Raper, J., Zimmermann, J., & Hayden, B. (2023). OpenApePose, a database of annotated ape photographs for pose estimation. eLife, 12, RP86873. https://doi.org/10.7554/eLife.86873.3

Data used for testing

Some of the chimpanzee samples used for testing were from the open-source data from:

  • Wiltshire, C., Lewis‐Cheetham, J., Komedová, V., Matsuzawa, T., Graham, K. E., & Hobaiter, C. (2023). DeepWild: Application of the pose estimation tool DeepLabCut for behaviour tracking in wild chimpanzees and bonobos. Journal of Animal Ecology, 92(8), 1560–1574. https://doi.org/10.1111/1365-2656.13932

Siamang samples were from YouTube and recordings related to the following work:

  • Pouw, W., Kehy, M., Gamba, M., & Ravignani, A. (2026). Amplitude increases of vocalizations are associated with body accelerations in siamang (Symphalangus syndactylus). International Journal of Comparative Psychology, 39. https://doi.org/10.46867/ijcp.53165

Demo

smoothing draw

Installation

GPU (NVIDIA):

conda create -n apetrack python=3.11 -y
conda activate apetrack
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu126
pip install envisionboxbio-apetrack

CPU only:

conda create -n apetrack python=3.11 -y
conda activate apetrack
pip install torch torchvision --index-url https://download.pytorch.org/whl/cpu
pip install envisionboxbio-apetrack

Note that the first time you run the tracker, it will download the OWLv2 detector and OpenApePose models (≈ 1.5 GB total), which will take some time. Make sure for GPU you use your CUDA version (versions are backward compatible). If you just want to try it out, just use cpu.

Tracking apes

from envisionboxbio_apetrack import ApeTracker

tracker = ApeTracker()                                   # default settings
tracker.process_video("clip.mp4", "output/")             # one video
tracker.process_folder("videos/", "output/")             # every video in a folder

# other settings, e.g.
tracker = ApeTracker(box_smoothing="medium", keypoint_smoothing="medium", draw=("skeleton", "points"))

Per video: <name>_tracks.csv (one row per ape per frame: box, track ID, 16 keypoints with likelihood), <name>_points.mp4 (default; <name>_skeleton.mp4 with draw=("skeleton",)), and <name>_detections.csv (raw detections, reused when you re-run with other tracking or smoothing settings).

Example output: the first rows of a _tracks.csv (two siamangs, default settings):

Gibbons_ChibaZoologicalPark_12_brachiating_vocalizing_tracks.csv: 195 rows (one per ape per frame) × 61 columns (download the full file). Columns: frame, time_s, track_id, det_conf, interpolated; box box_x/y/w/h (smoothed) and raw_box_x/y/w/h; then <keypoint>_x, _y, _likelihood for the 16 keypoints.

frame time_s track_id det_conf interpolated box_x box_y box_w box_h raw_box_x raw_box_y raw_box_w raw_box_h nose_x nose_y nose_likelihood left_eye_x left_eye_y left_eye_likelihood right_eye_x right_eye_y right_eye_likelihood head_x head_y head_likelihood neck_x neck_y neck_likelihood left_shoulder_x left_shoulder_y left_shoulder_likelihood left_elbow_x left_elbow_y left_elbow_likelihood left_wrist_x left_wrist_y left_wrist_likelihood right_shoulder_x right_shoulder_y right_shoulder_likelihood right_elbow_x right_elbow_y right_elbow_likelihood right_wrist_x right_wrist_y right_wrist_likelihood hip_x hip_y hip_likelihood left_knee_x left_knee_y left_knee_likelihood left_ankle_x left_ankle_y left_ankle_likelihood right_knee_x right_knee_y right_knee_likelihood right_ankle_x right_ankle_y right_ankle_likelihood
0 0.00 1 0.44 False 282.77 356.74 218.00 188.51 282.77 358.77 215.86 189.96 454.48 453.08 0.37 455.31 438.29 0.35 435.07 438.55 0.39 475.63 433.33 0.28 427.38 420.21 0.26 419.39 380.66 0.21 425.89 472.75 0.20 475.60 505.91 0.28 412.32 412.57 0.17 405.96 463.84 0.16 467.34 485.17 0.22 328.86 359.63 0.15 451.34 445.47 0.19 465.68 493.46 0.20 459.38 459.21 0.17 334.11 442.15 0.21
0 0.00 2 0.59 False 1357.90 497.47 199.11 346.13 1357.79 497.46 198.16 346.64 1419.70 675.76 0.42 1423.78 666.10 0.46 1422.79 660.16 0.39 1421.22 638.92 0.37 1459.91 656.38 0.43 1450.40 656.77 0.51 1395.18 624.61 0.64 1377.92 509.36 0.57 1490.01 656.09 0.33 1395.18 624.21 0.23 1374.83 501.72 0.26 1520.67 754.07 0.28 1443.54 746.22 0.68 1512.23 823.07 0.49 1458.09 748.92 0.19 1520.38 826.36 0.35
1 0.02 1 0.28 False 283.51 364.99 215.58 162.52 282.71 363.05 216.91 157.97 420.84 462.21 0.26 417.07 450.03 0.26 432.39 452.74 0.31 415.06 419.67 0.21 419.53 434.15 0.20 425.02 390.79 0.16 429.48 464.17 0.11 500.31 498.36 0.09 417.06 420.63 0.14 371.90 442.10 0.11 503.58 511.92 0.09 331.90 430.33 0.06 435.35 501.85 0.09 493.44 504.51 0.08 424.65 483.44 0.07 391.99 483.37 0.08
1 0.02 2 0.56 False 1368.32 496.60 197.19 351.11 1367.99 495.47 200.27 353.44 1435.13 675.57 0.37 1435.39 665.60 0.42 1440.13 660.91 0.34 1436.87 641.57 0.35 1471.75 657.70 0.44 1457.93 657.70 0.50 1406.08 624.27 0.65 1387.00 510.23 0.57 1503.11 657.70 0.32 1406.08 624.86 0.19 1385.55 501.72 0.23 1531.55 756.21 0.31 1452.38 746.82 0.67 1522.45 826.29 0.45 1474.96 747.67 0.18 1530.66 827.44 0.34
2 0.03 1 0.34 False 284.31 373.81 210.99 147.09 285.35 371.19 215.86 157.62 389.95 470.12 0.12 381.33 457.80 0.14 419.52 461.82 0.13 368.12 413.33 0.14 409.18 445.96 0.13 423.81 401.65 0.12 428.58 460.55 0.07 502.34 497.84 0.06 413.99 425.87 0.14 346.78 430.71 0.07 510.09 524.50 0.07 342.02 483.40 0.05 419.72 540.99 0.12 501.36 516.46 0.10 397.49 503.40 0.11 427.02 516.01 0.11
2 0.03 2 0.59 False 1378.46 496.48 195.93 354.16 1380.64 499.22 189.96 347.81 1449.51 675.78 0.32 1446.85 665.74 0.36 1455.73 661.91 0.29 1451.86 643.92 0.33 1483.14 658.97 0.43 1466.31 658.38 0.46 1416.45 624.14 0.66 1396.09 510.83 0.57 1515.61 659.26 0.29 1416.45 625.33 0.23 1395.80 501.97 0.23 1541.87 758.12 0.29 1461.59 747.20 0.67 1532.42 828.65 0.44 1489.31 746.28 0.23 1540.68 828.64 0.32
3 0.05 1 0.32 False 285.15 383.20 204.21 142.21 285.59 384.55 202.73 127.15 361.80 476.80 0.10 348.10 461.60 0.17 396.47 465.79 0.17 334.80 414.33 0.14 396.35 455.64 0.09 415.76 413.24 0.15 423.19 461.88 0.13 481.69 504.37 0.09 403.11 428.30 0.10 330.62 429.65 0.07 486.85 522.91 0.07 359.23 518.85 0.02 404.46 562.89 0.04 489.44 529.31 0.03 377.90 519.09 0.04 439.19 540.05 0.02
3 0.05 2 0.62 False 1388.32 497.12 195.32 355.27 1385.86 496.41 202.03 354.38 1462.85 676.40 0.40 1458.13 666.50 0.45 1469.59 663.15 0.36 1466.20 645.96 0.40 1494.10 660.19 0.46 1475.55 658.81 0.53 1426.32 624.22 0.69 1405.18 511.17 0.58 1527.51 660.78 0.36 1426.32 625.61 0.19 1405.56 502.47 0.23 1551.63 759.80 0.33 1471.19 747.35 0.68 1542.14 830.16 0.41 1501.15 744.76 0.17 1550.44 829.97 0.33

How videos are transformed

output
frame rate same as the input (e.g. 25 fps in → 25 fps out); one output frame and one CSV frame index per input frame
resolution same as the input unless output_width is set (the demo uses 960); CSV coordinates always in input pixels
interlaced video deinterlaced, one frame per frame (frame rate unchanged)
audio copied into the labelled videos
codec H.264 .mp4

Settings

setting default
detector_model ‘google/owlv2-base-patch16-ensemble’
prompts (‘a photo of a chimpanzee’, ‘a photo of a bonobo’, ‘a photo of a gorilla’, ‘a photo of an orangutan’, ‘a photo of a gibbon’, ‘a photo of a siamang’, ‘a photo of an ape’)
detection_threshold 0.15
nms_iou 0.7
tiles 1
detect_every 1
track_activation_threshold 0.25
track_buffer_s 1.0
track_matching_threshold 0.9
stitch_gap_s 2.0
stitch_distance 1.0
fill_gap_s 0.5
box_smoothing ‘low’
keypoint_smoothing ‘low’
pad 0.1
pose_weights None
draw (‘points’,)
output_width None
min_likelihood_draw 0.0
device None

Smoothing windows (seconds, Savitzky-Golay order 2): none 0.0, low 0.1, medium 0.25, high 0.5

Under the hood

The whole pipeline (ApeTracker.process_video)
    def process_video(self, video_path, output_folder, reuse_detections=True):
        """Writes <name>_detections.csv (raw detections, reused next time), <name>_tracks.csv and
        <name>_<style>.mp4 per drawing style. Returns the tracks table."""
        c = self.config
        os.makedirs(output_folder, exist_ok=True)
        name = os.path.splitext(os.path.basename(video_path))[0]
        det_path = os.path.join(output_folder, f"{name}_detections.csv")

        # 1. detect (slow; saved)
        if reuse_detections and os.path.exists(det_path) and _same_detection_settings(det_path, c):
            dets = pd.read_csv(det_path)
            fps = float(dets.fps.iloc[0]) if len(dets) else video.read_video(video_path)[0]
        else:
            fps, size, frames = video.read_video(video_path)
            with tqdm(desc=f"{name}: detecting", unit="frame", leave=False) as bar:
                dets = detect.detect_video(self.detector, frames, fps, every=c.detect_every,
                                           threshold=min(0.05, c.detection_threshold), progress=bar)
            dets.assign(tiles=c.tiles).round(3).to_csv(det_path, index=False)

        # 2. track, stitch, fill gaps, smooth boxes
        tracks = track.bytetrack(dets, fps, c.detection_threshold, c.nms_iou, c.track_activation_threshold,
                                 c.track_buffer_s, c.track_matching_threshold)
        tracks = track.stitch_tracks(tracks, fps, c.stitch_gap_s, c.stitch_distance)
        tracks = track.fill_gaps(tracks, fps, max(c.fill_gap_s, c.stitch_gap_s, c.detect_every / fps))
        tracks = smooth.smooth_boxes(tracks, fps, c.box_smoothing)

        # 3. keypoints in every box, then smooth them
        kp = np.zeros((len(tracks), len(pose.KEYPOINTS), 3))
        rows_of = tracks.groupby("frame").indices
        fps, size, frames = video.read_video(video_path)
        for i, frame in enumerate(tqdm(frames, desc=f"{name}: pose", unit="frame", leave=False)):
            idx = rows_of.get(i)
            if idx is not None:
                kp[idx] = pose.keypoints_for_boxes(self.pose_model, frame, tracks.iloc[idx][track.BOX_COLS].to_numpy(),
                                                   c.pad, self.device)
        for k, name_k in enumerate(pose.KEYPOINTS):
            tracks[f"{name_k}_x"], tracks[f"{name_k}_y"], tracks[f"{name_k}_likelihood"] = kp[:, k, 0], kp[:, k, 1], kp[:, k, 2]
        tracks = smooth.smooth_keypoints(tracks, pose.KEYPOINTS, fps, c.keypoint_smoothing)
        tracks = tracks.sort_values(["frame", "track_id"]).reset_index(drop=True)
        tracks.round(3).to_csv(os.path.join(output_folder, f"{name}_tracks.csv"), index=False)

        # 4. labelled video(s): same frame rate as the input, input audio copied
        if c.draw:
            self._render(video_path, output_folder, name, tracks)
        return tracks
Models: OWLv2 detector (prompts, fp16, tiles)
class OWLv2Detector:
    def __init__(self, model_name, prompts, device, tiles=1, overlap=0.25):
        from transformers import Owlv2ForObjectDetection, Owlv2Processor
        try:                                                         # cached: no network access at all
            self.proc = Owlv2Processor.from_pretrained(model_name, local_files_only=True)
            self.model = Owlv2ForObjectDetection.from_pretrained(model_name, local_files_only=True)
        except OSError:                                              # first use: one-time download (~0.6 GB)
            self.proc = Owlv2Processor.from_pretrained(model_name)
            self.model = Owlv2ForObjectDetection.from_pretrained(model_name)
        self.model = self.model.to(device).eval()
        self.prompts = [list(prompts)]
        self.device = device
        self.fp16 = device.startswith("cuda")                        # ~5x faster, same boxes (within 2 px)
        self.tiles, self.overlap = tiles, overlap

    def __call__(self, frame_bgr, threshold):
        """All boxes scoring >= threshold, overlapping boxes NOT merged: (xyxy array, scores)."""
        if self.tiles <= 1:
            return self._detect(frame_bgr, threshold)
        H, W = frame_bgr.shape[:2]
        n, ov = self.tiles, self.overlap
        tw, th = W / (n - (n - 1) * ov), H / (n - (n - 1) * ov)
        b0, s0 = self._detect(frame_bgr, threshold)
        boxes, scores = [b0], [s0]
        for iy in range(n):
            for ix in range(n):
                x0, y0 = int(ix * tw * (1 - ov)), int(iy * th * (1 - ov))
                b, s = self._detect(frame_bgr[y0:int(y0 + th), x0:int(x0 + tw)], threshold)
                boxes.append(b + [x0, y0, x0, y0])
                scores.append(s)
        return np.concatenate(boxes), np.concatenate(scores)

    def _detect(self, frame_bgr, threshold):
        im = Image.fromarray(cv2.cvtColor(frame_bgr, cv2.COLOR_BGR2RGB))
        inp = self.proc(text=self.prompts, images=im, return_tensors="pt").to(self.device)
        side = max(im.size)                                          # OWLv2 pads the image to a square
        with torch.no_grad(), torch.autocast("cuda", dtype=torch.float16, enabled=self.fp16):
            out = self.model(**inp)
        out.logits, out.pred_boxes = out.logits.float(), out.pred_boxes.float()
        r = self.proc.post_process_grounded_object_detection(out, threshold=threshold, target_sizes=[(side, side)])[0]
        xyxy = np.clip(r["boxes"].cpu().numpy().astype(float), 0, [im.size[0], im.size[1]] * 2)
        return xyxy, r["scores"].cpu().numpy().astype(float)
Models: OpenApePose (HRNet-W48, weights, flip test)
class OpenApePose(torch.nn.Module):
    def __init__(self):
        super().__init__()
        import timm
        self.backbone = timm.create_model("hrnet_w48", pretrained=False, features_only=True, feature_location="")
        self.head = torch.nn.Conv2d(48, len(KEYPOINTS), 1)

    def forward(self, x):
        return self.head(self.backbone(x)[1])                        # [1] = 48-channel branch at 1/4 resolution


def load_model(weights_path, device):
    """The fp16 safetensors file holds the complete state dict of OpenApePose(): strict loading."""
    from safetensors.torch import load_file
    model = OpenApePose()
    model.load_state_dict({k: v.float() for k, v in load_file(weights_path).items()}, strict=True)
    return model.to(device).eval()


def keypoints_for_boxes(model, frame_bgr, boxes_xywh, pad, device):
    """All boxes of one frame in one batch (plus mirrored crops). Returns (n_boxes, 16, 3): x, y, likelihood."""
    if not len(boxes_xywh):
        return np.zeros((0, len(KEYPOINTS), 3))
    crops, Ms = [], []
    for x, y, w, h in boxes_xywh:
        M = crop_matrix(x - pad * w, y - pad * h, w * (1 + 2 * pad), h * (1 + 2 * pad))
        c = (cv2.warpAffine(frame_bgr, M, (W, H))[:, :, ::-1] / 255.0 - MEAN) / STD       # BGR -> RGB
        crops += [c, c[:, ::-1]]
        Ms.append(M)
    batch = torch.tensor(np.stack(crops).transpose(0, 3, 1, 2).copy(), dtype=torch.float32, device=device)
    with torch.no_grad():
        hm = model(batch).cpu().numpy()
    out = []
    for b, M in enumerate(Ms):
        mirrored = hm[2 * b + 1][FLIP][:, :, ::-1]                   # undo mirroring, swap left/right
        mirrored = np.concatenate([mirrored[:, :, :1], mirrored[:, :, :-1]], axis=2)      # 1-pixel shift
        px, py, conf = peaks((hm[2 * b] + mirrored) / 2)
        back = cv2.invertAffineTransform(M)
        out.append(np.c_[back[0, 0] * px + back[0, 1] * py + back[0, 2], back[1, 0] * px + back[1, 1] * py + back[1, 2], conf])
    return np.array(out)
Settings (Config)
"""Default settings. Detection/tracking defaults were tuned on hand-checked siamang boxes (11 clips) and
DeepWild (Wiltshire et al. 2023; 1577 hand-labelled chimpanzees and bonobos)."""
from dataclasses import dataclass, field
from typing import Optional, Tuple

# Savitzky-Golay window in seconds (polynomial order 2). "high" lags fast movements such as brachiation.
SMOOTHING = {"none": 0.0, "low": 0.1, "medium": 0.25, "high": 0.5}

POSE_WEIGHTS_FILE = "openapepose_hrnet_w48_fp16.safetensors"
POSE_WEIGHTS_URL = "https://zenodo.org/records/22989835/files/openapepose_hrnet_w48_fp16.safetensors?download=1"
POSE_WEIGHTS_SHA256 = "9ac7636c3b086367a9728164537642c5d87020f1337d181e4ba779330afea8c7"


@dataclass
class Config:
    # --- detection: OWLv2 (Minderer et al. 2023), open-vocabulary, prompted with text ---
    detector_model: str = "google/owlv2-base-patch16-ensemble"
    prompts: Tuple[str, ...] = ("a photo of a chimpanzee", "a photo of a bonobo", "a photo of a gorilla",
                                "a photo of an orangutan", "a photo of a gibbon", "a photo of a siamang",
                                "a photo of an ape")
    detection_threshold: float = 0.15       # keep boxes scoring at least this
    nms_iou: float = 0.7                    # merge boxes that overlap more than this
    tiles: int = 1                          # >1: also detect on tiles x tiles enlarged crops (far-away apes; slower)
    detect_every: int = 1                   # detect on every n-th frame only (CPU); boxes in between are interpolated

    # --- tracking: ByteTrack (Zhang et al. 2022, via supervision) + stitching + gap filling ---
    track_activation_threshold: float = 0.25   # score needed to start a track
    track_buffer_s: float = 1.0                # how long a lost track is kept
    track_matching_threshold: float = 0.9      # 1 - IoU needed to link a box to a track (lenient: fast apes)
    stitch_gap_s: float = 2.0                  # re-join a track that reappears within this time (occlusion)
    stitch_distance: float = 1.0               # ... within this many ape sizes of where it disappeared
    fill_gap_s: float = 0.5                    # interpolate missing boxes within a track up to this long

    # --- smoothing ---
    box_smoothing: str = "low"                 # none | low | medium | high
    keypoint_smoothing: str = "low"            # none | low | medium | high

    # --- pose: OpenApePose (Desai et al. 2023) ---
    pad: float = 0.1                           # context around the box given to the pose model
    pose_weights: Optional[str] = None         # local .safetensors file; None = download on first use

    # --- output video ---
    draw: Tuple[str, ...] = ("points",)        # any of "points", "skeleton"; () = no video
    output_width: Optional[int] = None         # None = same size as the input
    min_likelihood_draw: float = 0.0           # hide keypoints/limbs below this likelihood
    device: Optional[str] = None               # None = cuda if available, else cpu

    extra: dict = field(default_factory=dict)
Tracking: ByteTrack, stitching, gap filling
def bytetrack(dets, fps, threshold, nms_iou, activation, buffer_s, matching):
    every = int(dets.every.iloc[0]) if "every" in dets and len(dets) else 1
    n_frames = int(dets.n_frames.iloc[0]) if len(dets) else 0
    rate = fps / every
    tr = sv.ByteTrack(track_activation_threshold=activation, lost_track_buffer=max(1, int(round(buffer_s * rate))),
                      minimum_matching_threshold=matching, frame_rate=max(1, int(round(rate))))
    tr.reset()                                                       # track IDs start at 1 in every video
    by_frame = {f: g for f, g in dets[dets.conf >= threshold].groupby("frame")}
    rows = []
    for i in range(0, n_frames, every):
        g = by_frame.get(i)
        d = sv.Detections(xyxy=g[["x0", "y0", "x1", "y1"]].to_numpy(float), confidence=g.conf.to_numpy(float),
                          class_id=np.zeros(len(g), int)) if g is not None else sv.Detections.empty()
        if len(d):
            d = d.with_nms(threshold=nms_iou, class_agnostic=True)
        d = tr.update_with_detections(d)
        for (x0, y0, x1, y1), tid, c in zip(d.xyxy, d.tracker_id, d.confidence):
            rows.append({"frame": i, "time_s": i / fps, "track_id": int(tid), "det_conf": float(c), "interpolated": False,
                         "box_x": x0, "box_y": y0, "box_w": x1 - x0, "box_h": y1 - y0})
    return pd.DataFrame(rows, columns=TRACK_COLS)


def stitch_tracks(boxes, fps, max_gap_s, max_dist=1.0, max_size_ratio=2.0):
    """A track that starts 1 frame..max_gap_s after another ended, within max_dist x ape size (sqrt box area) of
    where that one ended and at a similar size, gets the older track's ID (an ape that was hidden and
    reappeared). Joins are one-to-one, closest first, so two apes cannot take over each other's ID."""
    if max_gap_s <= 0 or not len(boxes):
        return boxes
    b = boxes.sort_values("frame")
    ends = b.groupby("track_id").tail(1).set_index("track_id")
    starts = b.groupby("track_id").head(1).set_index("track_id")
    cand = []
    for old, e in ends.iterrows():
        es = np.sqrt(e.box_w * e.box_h)
        for new, st in starts.iterrows():
            gap = st.frame - e.frame
            if new == old or gap < 1 or gap > max_gap_s * fps:
                continue
            ss = np.sqrt(st.box_w * st.box_h)
            dist = np.hypot(st.box_x + st.box_w / 2 - e.box_x - e.box_w / 2, st.box_y + st.box_h / 2 - e.box_y - e.box_h / 2)
            if dist < max_dist * es and max(ss, es) / min(ss, es) < max_size_ratio:
                cand.append((dist / es, old, new))
    joined, used_old, used_new = {}, set(), set()
    for _, old, new in sorted(cand):
        if old not in used_old and new not in used_new:
            joined[new] = old
            used_old.add(old)
            used_new.add(new)

    def root(t):
        while t in joined:
            t = joined[t]
        return t
    return boxes.assign(track_id=boxes.track_id.map(root))


def fill_gaps(boxes, fps, max_gap_s):
    """Linearly interpolate a track's box over gaps up to max_gap_s (rows get interpolated=True)."""
    max_gap = int(round(max_gap_s * fps))
    if max_gap < 1 or not len(boxes):
        return boxes
    new = []
    for tid, g in boxes.sort_values("frame").groupby("track_id"):
        f = g.frame.to_numpy()
        for a in np.flatnonzero(np.diff(f) > 1):
            if f[a + 1] - f[a] - 1 > max_gap:
                continue
            ra, rb = g.iloc[a], g.iloc[a + 1]
            for fr in range(f[a] + 1, f[a + 1]):
                t = (fr - f[a]) / (f[a + 1] - f[a])
                row = {c: ra[c] + t * (rb[c] - ra[c]) for c in BOX_COLS}
                row.update(frame=fr, time_s=fr / fps, track_id=tid, det_conf=np.nan, interpolated=True)
                new.append(row)
    out = pd.concat([boxes, pd.DataFrame(new, columns=TRACK_COLS)], ignore_index=True) if new else boxes
    return out.sort_values(["frame", "track_id"]).reset_index(drop=True)
Smoothing: boxes and keypoints
"""Savitzky-Golay smoothing (order 2) per track, over runs of consecutive frames only (never across frames where
the ape was not tracked). Window = SMOOTHING[level] seconds, at least 5 frames (order 2 needs 5 to smooth)."""
import numpy as np
from scipy.signal import savgol_filter

from .config import SMOOTHING


def _window(level, fps):
    if SMOOTHING[level] == 0:
        return 0
    return max(int(round(SMOOTHING[level] * fps)) | 1, 5)


def _smooth_runs(df, cols, win):
    """Smooth columns `cols` of df (sorted by track, frame) within runs of consecutive frames."""
    sig = df[cols].to_numpy(float)
    run = ((df.frame.diff() != 1) | (df.track_id.diff() != 0)).cumsum().to_numpy()
    for r in np.unique(run):
        idx = np.flatnonzero(run == r)
        w = min(win, len(idx) if len(idx) % 2 else len(idx) - 1)
        if w >= 5:
            sig[idx] = savgol_filter(sig[idx], w, 2, axis=0)
    return sig


def smooth_boxes(boxes, fps, level):
    """Smooths each box's centre, width and height. Keeps the unsmoothed box as raw_box_*."""
    out = boxes.sort_values(["track_id", "frame"]).reset_index(drop=True)
    for c in ["box_x", "box_y", "box_w", "box_h"]:
        out["raw_" + c] = out[c]
    win = _window(level, fps)
    if not win or not len(out):
        return out
    out["cx"], out["cy"] = out.box_x + out.box_w / 2, out.box_y + out.box_h / 2
    s = _smooth_runs(out, ["cx", "cy", "box_w", "box_h"], win)
    out["box_w"], out["box_h"] = np.maximum(s[:, 2], 1), np.maximum(s[:, 3], 1)
    out["box_x"], out["box_y"] = s[:, 0] - out.box_w / 2, s[:, 1] - out.box_h / 2
    return out.drop(columns=["cx", "cy"])


def smooth_keypoints(tracks, keypoints, fps, level):
    """Smooths every keypoint's x and y over time (likelihoods are left as they are)."""
    win = _window(level, fps)
    if not win or not len(tracks):
        return tracks
    out = tracks.sort_values(["track_id", "frame"]).reset_index(drop=True)
    cols = [f"{k}_{a}" for k in keypoints for a in "xy"]
    out[cols] = _smooth_runs(out, cols, win)
    return out

Limitations

Full evaluation still needs to be performed. Distant apes or overlapping apes are not tracked well. It is limited to apes, monkeys will likely not be tracked well.

Evaluation

  • Planned (if you want to collaborate on this reach out to w.pouw@tilburguniversity.edu)

OpenApe Pose model

We redistributed the trained OpenApePose model (HRNet-W48, file hrnet_w48_oap_256x192_full.pth) on Zenodo (https://zenodo.org/records/22989835), which was originally deposited on Dryad under public domain (CC0) (https://doi.org/10.5061/dryad.c59zw3rds) and here released under the MIT license in the repository (https://github.com/desai-nisarg/OpenApePose). The only change is a conversion to safetensors [fp16], keeping the backbone and keypoint head weights; the model was not retrained or modified. Please cite the original authors (Desai et al., 2023).

Licence

MIT. Model weights: OWLv2 Apache-2.0; OpenApePose weights CC0 license.

Full documentation (installation, usage, outputs, settings, speed): see the README.