Pipeline: OWLv2 (boxes, prompted with ape names) -> ByteTrack (track IDs within a video) + stitching across short occlusions -> Savitzky-Golay smoothing bounding boxes -> OpenApePose (16 keypoints per ape) -> Savitzky-Golay smoothing smoothing keypoints -> CSV + labelled video.
The current tool serves as an easy to use “out of the box” pose tracker for apes. While there are many easy to use deep learning frameworks to train one’s own model with your own labeled data, to my surprise there were no easy to use ape-specific tracking that work on any video. The available software, like OpenApePose (Desai et al., 2023) for keypoint detection, were however there to be easily combined with powerful top-down bounding box detection models (Owlv2), which together with software for stable detection over sequences (ByteTrack), makes a promising pose estimation tool.
OpenApePose (Desai, 2023) was trained on 71,868 annotated images:
18,010 chimpanzees (Pan troglodytes)
11,685 bonobos (Pan paniscus)
12,905 gorillas (Gorilla gorilla)
12,722 orangutans (Pongo sp.)
9274 gibbons (genus Hylobates and Nomascus)
7272 siamangs (Symphalangus syndactylus)
Please reach out to w.pouw@tilburguniversity.edu if you want to help out with proper evaluation against state of the art (but usually less user-friendly) computer vision models.
Citation
Citing envisionboxbio-apetrack
Pouw, W., Desai, N. (2026). envisionboxbio-apetrack: Zero-shot multi-ape tracking and pose estimation using OWLv2, ByteTrack and OpenApePose (Version 0.1.1) [Computer software]. Zenodo. https://doi.org/10.5281/zenodo.22992133
This python package is a productive recombination of available and openly licensed software. Therefore first and foremost these packages need to be cited:
OWLv2: Minderer, M., Gritsenko, A., & Houlsby, N. (2023). Scaling open-vocabulary object detection. Advances in Neural Information Processing Systems, 36, 72983–73007. https://doi.org/10.52202/075280-3191
ByteTrack: Zhang, Y., Sun, P., Jiang, Y., Yu, D., Weng, F., Yuan, Z., … & Wang, X. (2022). ByteTrack: Multi-object tracking by associating every detection box. In European Conference on Computer Vision (pp. 1–21). Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-20047-2_1
OpenApePose: Desai, N., Bala, P., Richardson, R., Raper, J., Zimmermann, J., & Hayden, B. (2023). OpenApePose, a database of annotated ape photographs for pose estimation. eLife, 12, RP86873. https://doi.org/10.7554/eLife.86873.3
Data used for testing
Some of the chimpanzee samples used for testing were from the open-source data from:
Wiltshire, C., Lewis‐Cheetham, J., Komedová, V., Matsuzawa, T., Graham, K. E., & Hobaiter, C. (2023). DeepWild: Application of the pose estimation tool DeepLabCut for behaviour tracking in wild chimpanzees and bonobos. Journal of Animal Ecology, 92(8), 1560–1574. https://doi.org/10.1111/1365-2656.13932
Siamang samples were from YouTube and recordings related to the following work:
Pouw, W., Kehy, M., Gamba, M., & Ravignani, A. (2026). Amplitude increases of vocalizations are associated with body accelerations in siamang (Symphalangus syndactylus). International Journal of Comparative Psychology, 39. https://doi.org/10.46867/ijcp.53165
Note that the first time you run the tracker, it will download the OWLv2 detector and OpenApePose models (≈ 1.5 GB total), which will take some time. Make sure for GPU you use your CUDA version (versions are backward compatible). If you just want to try it out, just use cpu.
Tracking apes
from envisionboxbio_apetrack import ApeTrackertracker = ApeTracker() # default settingstracker.process_video("clip.mp4", "output/") # one videotracker.process_folder("videos/", "output/") # every video in a folder# other settings, e.g.tracker = ApeTracker(box_smoothing="medium", keypoint_smoothing="medium", draw=("skeleton", "points"))
Per video: <name>_tracks.csv (one row per ape per frame: box, track ID, 16 keypoints with likelihood), <name>_points.mp4 (default; <name>_skeleton.mp4 with draw=("skeleton",)), and <name>_detections.csv (raw detections, reused when you re-run with other tracking or smoothing settings).
Example output: the first rows of a _tracks.csv (two siamangs, default settings):
Gibbons_ChibaZoologicalPark_12_brachiating_vocalizing_tracks.csv: 195 rows (one per ape per frame) × 61 columns (download the full file). Columns: frame, time_s, track_id, det_conf, interpolated; box box_x/y/w/h (smoothed) and raw_box_x/y/w/h; then <keypoint>_x, _y, _likelihood for the 16 keypoints.
frame
time_s
track_id
det_conf
interpolated
box_x
box_y
box_w
box_h
raw_box_x
raw_box_y
raw_box_w
raw_box_h
nose_x
nose_y
nose_likelihood
left_eye_x
left_eye_y
left_eye_likelihood
right_eye_x
right_eye_y
right_eye_likelihood
head_x
head_y
head_likelihood
neck_x
neck_y
neck_likelihood
left_shoulder_x
left_shoulder_y
left_shoulder_likelihood
left_elbow_x
left_elbow_y
left_elbow_likelihood
left_wrist_x
left_wrist_y
left_wrist_likelihood
right_shoulder_x
right_shoulder_y
right_shoulder_likelihood
right_elbow_x
right_elbow_y
right_elbow_likelihood
right_wrist_x
right_wrist_y
right_wrist_likelihood
hip_x
hip_y
hip_likelihood
left_knee_x
left_knee_y
left_knee_likelihood
left_ankle_x
left_ankle_y
left_ankle_likelihood
right_knee_x
right_knee_y
right_knee_likelihood
right_ankle_x
right_ankle_y
right_ankle_likelihood
0
0.00
1
0.44
False
282.77
356.74
218.00
188.51
282.77
358.77
215.86
189.96
454.48
453.08
0.37
455.31
438.29
0.35
435.07
438.55
0.39
475.63
433.33
0.28
427.38
420.21
0.26
419.39
380.66
0.21
425.89
472.75
0.20
475.60
505.91
0.28
412.32
412.57
0.17
405.96
463.84
0.16
467.34
485.17
0.22
328.86
359.63
0.15
451.34
445.47
0.19
465.68
493.46
0.20
459.38
459.21
0.17
334.11
442.15
0.21
0
0.00
2
0.59
False
1357.90
497.47
199.11
346.13
1357.79
497.46
198.16
346.64
1419.70
675.76
0.42
1423.78
666.10
0.46
1422.79
660.16
0.39
1421.22
638.92
0.37
1459.91
656.38
0.43
1450.40
656.77
0.51
1395.18
624.61
0.64
1377.92
509.36
0.57
1490.01
656.09
0.33
1395.18
624.21
0.23
1374.83
501.72
0.26
1520.67
754.07
0.28
1443.54
746.22
0.68
1512.23
823.07
0.49
1458.09
748.92
0.19
1520.38
826.36
0.35
1
0.02
1
0.28
False
283.51
364.99
215.58
162.52
282.71
363.05
216.91
157.97
420.84
462.21
0.26
417.07
450.03
0.26
432.39
452.74
0.31
415.06
419.67
0.21
419.53
434.15
0.20
425.02
390.79
0.16
429.48
464.17
0.11
500.31
498.36
0.09
417.06
420.63
0.14
371.90
442.10
0.11
503.58
511.92
0.09
331.90
430.33
0.06
435.35
501.85
0.09
493.44
504.51
0.08
424.65
483.44
0.07
391.99
483.37
0.08
1
0.02
2
0.56
False
1368.32
496.60
197.19
351.11
1367.99
495.47
200.27
353.44
1435.13
675.57
0.37
1435.39
665.60
0.42
1440.13
660.91
0.34
1436.87
641.57
0.35
1471.75
657.70
0.44
1457.93
657.70
0.50
1406.08
624.27
0.65
1387.00
510.23
0.57
1503.11
657.70
0.32
1406.08
624.86
0.19
1385.55
501.72
0.23
1531.55
756.21
0.31
1452.38
746.82
0.67
1522.45
826.29
0.45
1474.96
747.67
0.18
1530.66
827.44
0.34
2
0.03
1
0.34
False
284.31
373.81
210.99
147.09
285.35
371.19
215.86
157.62
389.95
470.12
0.12
381.33
457.80
0.14
419.52
461.82
0.13
368.12
413.33
0.14
409.18
445.96
0.13
423.81
401.65
0.12
428.58
460.55
0.07
502.34
497.84
0.06
413.99
425.87
0.14
346.78
430.71
0.07
510.09
524.50
0.07
342.02
483.40
0.05
419.72
540.99
0.12
501.36
516.46
0.10
397.49
503.40
0.11
427.02
516.01
0.11
2
0.03
2
0.59
False
1378.46
496.48
195.93
354.16
1380.64
499.22
189.96
347.81
1449.51
675.78
0.32
1446.85
665.74
0.36
1455.73
661.91
0.29
1451.86
643.92
0.33
1483.14
658.97
0.43
1466.31
658.38
0.46
1416.45
624.14
0.66
1396.09
510.83
0.57
1515.61
659.26
0.29
1416.45
625.33
0.23
1395.80
501.97
0.23
1541.87
758.12
0.29
1461.59
747.20
0.67
1532.42
828.65
0.44
1489.31
746.28
0.23
1540.68
828.64
0.32
3
0.05
1
0.32
False
285.15
383.20
204.21
142.21
285.59
384.55
202.73
127.15
361.80
476.80
0.10
348.10
461.60
0.17
396.47
465.79
0.17
334.80
414.33
0.14
396.35
455.64
0.09
415.76
413.24
0.15
423.19
461.88
0.13
481.69
504.37
0.09
403.11
428.30
0.10
330.62
429.65
0.07
486.85
522.91
0.07
359.23
518.85
0.02
404.46
562.89
0.04
489.44
529.31
0.03
377.90
519.09
0.04
439.19
540.05
0.02
3
0.05
2
0.62
False
1388.32
497.12
195.32
355.27
1385.86
496.41
202.03
354.38
1462.85
676.40
0.40
1458.13
666.50
0.45
1469.59
663.15
0.36
1466.20
645.96
0.40
1494.10
660.19
0.46
1475.55
658.81
0.53
1426.32
624.22
0.69
1405.18
511.17
0.58
1527.51
660.78
0.36
1426.32
625.61
0.19
1405.56
502.47
0.23
1551.63
759.80
0.33
1471.19
747.35
0.68
1542.14
830.16
0.41
1501.15
744.76
0.17
1550.44
829.97
0.33
How videos are transformed
output
frame rate
same as the input (e.g. 25 fps in → 25 fps out); one output frame and one CSV frame index per input frame
resolution
same as the input unless output_width is set (the demo uses 960); CSV coordinates always in input pixels
interlaced video
deinterlaced, one frame per frame (frame rate unchanged)
audio
copied into the labelled videos
codec
H.264 .mp4
Settings
setting
default
detector_model
‘google/owlv2-base-patch16-ensemble’
prompts
(‘a photo of a chimpanzee’, ‘a photo of a bonobo’, ‘a photo of a gorilla’, ‘a photo of an orangutan’, ‘a photo of a gibbon’, ‘a photo of a siamang’, ‘a photo of an ape’)
detection_threshold
0.15
nms_iou
0.7
tiles
1
detect_every
1
track_activation_threshold
0.25
track_buffer_s
1.0
track_matching_threshold
0.9
stitch_gap_s
2.0
stitch_distance
1.0
fill_gap_s
0.5
box_smoothing
‘low’
keypoint_smoothing
‘low’
pad
0.1
pose_weights
None
draw
(‘points’,)
output_width
None
min_likelihood_draw
0.0
device
None
Smoothing windows (seconds, Savitzky-Golay order 2): none 0.0, low 0.1, medium 0.25, high 0.5
Under the hood
The whole pipeline (ApeTracker.process_video)
def process_video(self, video_path, output_folder, reuse_detections=True):"""Writes <name>_detections.csv (raw detections, reused next time), <name>_tracks.csv and <name>_<style>.mp4 per drawing style. Returns the tracks table.""" c =self.config os.makedirs(output_folder, exist_ok=True) name = os.path.splitext(os.path.basename(video_path))[0] det_path = os.path.join(output_folder, f"{name}_detections.csv")# 1. detect (slow; saved)if reuse_detections and os.path.exists(det_path) and _same_detection_settings(det_path, c): dets = pd.read_csv(det_path) fps =float(dets.fps.iloc[0]) iflen(dets) else video.read_video(video_path)[0]else: fps, size, frames = video.read_video(video_path)with tqdm(desc=f"{name}: detecting", unit="frame", leave=False) as bar: dets = detect.detect_video(self.detector, frames, fps, every=c.detect_every, threshold=min(0.05, c.detection_threshold), progress=bar) dets.assign(tiles=c.tiles).round(3).to_csv(det_path, index=False)# 2. track, stitch, fill gaps, smooth boxes tracks = track.bytetrack(dets, fps, c.detection_threshold, c.nms_iou, c.track_activation_threshold, c.track_buffer_s, c.track_matching_threshold) tracks = track.stitch_tracks(tracks, fps, c.stitch_gap_s, c.stitch_distance) tracks = track.fill_gaps(tracks, fps, max(c.fill_gap_s, c.stitch_gap_s, c.detect_every / fps)) tracks = smooth.smooth_boxes(tracks, fps, c.box_smoothing)# 3. keypoints in every box, then smooth them kp = np.zeros((len(tracks), len(pose.KEYPOINTS), 3)) rows_of = tracks.groupby("frame").indices fps, size, frames = video.read_video(video_path)for i, frame inenumerate(tqdm(frames, desc=f"{name}: pose", unit="frame", leave=False)): idx = rows_of.get(i)if idx isnotNone: kp[idx] = pose.keypoints_for_boxes(self.pose_model, frame, tracks.iloc[idx][track.BOX_COLS].to_numpy(), c.pad, self.device)for k, name_k inenumerate(pose.KEYPOINTS): tracks[f"{name_k}_x"], tracks[f"{name_k}_y"], tracks[f"{name_k}_likelihood"] = kp[:, k, 0], kp[:, k, 1], kp[:, k, 2] tracks = smooth.smooth_keypoints(tracks, pose.KEYPOINTS, fps, c.keypoint_smoothing) tracks = tracks.sort_values(["frame", "track_id"]).reset_index(drop=True) tracks.round(3).to_csv(os.path.join(output_folder, f"{name}_tracks.csv"), index=False)# 4. labelled video(s): same frame rate as the input, input audio copiedif c.draw:self._render(video_path, output_folder, name, tracks)return tracks
Models: OWLv2 detector (prompts, fp16, tiles)
class OWLv2Detector:def__init__(self, model_name, prompts, device, tiles=1, overlap=0.25):from transformers import Owlv2ForObjectDetection, Owlv2Processortry: # cached: no network access at allself.proc = Owlv2Processor.from_pretrained(model_name, local_files_only=True)self.model = Owlv2ForObjectDetection.from_pretrained(model_name, local_files_only=True)exceptOSError: # first use: one-time download (~0.6 GB)self.proc = Owlv2Processor.from_pretrained(model_name)self.model = Owlv2ForObjectDetection.from_pretrained(model_name)self.model =self.model.to(device).eval()self.prompts = [list(prompts)]self.device = deviceself.fp16 = device.startswith("cuda") # ~5x faster, same boxes (within 2 px)self.tiles, self.overlap = tiles, overlapdef__call__(self, frame_bgr, threshold):"""All boxes scoring >= threshold, overlapping boxes NOT merged: (xyxy array, scores)."""ifself.tiles <=1:returnself._detect(frame_bgr, threshold) H, W = frame_bgr.shape[:2] n, ov =self.tiles, self.overlap tw, th = W / (n - (n -1) * ov), H / (n - (n -1) * ov) b0, s0 =self._detect(frame_bgr, threshold) boxes, scores = [b0], [s0]for iy inrange(n):for ix inrange(n): x0, y0 =int(ix * tw * (1- ov)), int(iy * th * (1- ov)) b, s =self._detect(frame_bgr[y0:int(y0 + th), x0:int(x0 + tw)], threshold) boxes.append(b + [x0, y0, x0, y0]) scores.append(s)return np.concatenate(boxes), np.concatenate(scores)def _detect(self, frame_bgr, threshold): im = Image.fromarray(cv2.cvtColor(frame_bgr, cv2.COLOR_BGR2RGB)) inp =self.proc(text=self.prompts, images=im, return_tensors="pt").to(self.device) side =max(im.size) # OWLv2 pads the image to a squarewith torch.no_grad(), torch.autocast("cuda", dtype=torch.float16, enabled=self.fp16): out =self.model(**inp) out.logits, out.pred_boxes = out.logits.float(), out.pred_boxes.float() r =self.proc.post_process_grounded_object_detection(out, threshold=threshold, target_sizes=[(side, side)])[0] xyxy = np.clip(r["boxes"].cpu().numpy().astype(float), 0, [im.size[0], im.size[1]] *2)return xyxy, r["scores"].cpu().numpy().astype(float)
class OpenApePose(torch.nn.Module):def__init__(self):super().__init__()import timmself.backbone = timm.create_model("hrnet_w48", pretrained=False, features_only=True, feature_location="")self.head = torch.nn.Conv2d(48, len(KEYPOINTS), 1)def forward(self, x):returnself.head(self.backbone(x)[1]) # [1] = 48-channel branch at 1/4 resolutiondef load_model(weights_path, device):"""The fp16 safetensors file holds the complete state dict of OpenApePose(): strict loading."""from safetensors.torch import load_file model = OpenApePose() model.load_state_dict({k: v.float() for k, v in load_file(weights_path).items()}, strict=True)return model.to(device).eval()def keypoints_for_boxes(model, frame_bgr, boxes_xywh, pad, device):"""All boxes of one frame in one batch (plus mirrored crops). Returns (n_boxes, 16, 3): x, y, likelihood."""ifnotlen(boxes_xywh):return np.zeros((0, len(KEYPOINTS), 3)) crops, Ms = [], []for x, y, w, h in boxes_xywh: M = crop_matrix(x - pad * w, y - pad * h, w * (1+2* pad), h * (1+2* pad)) c = (cv2.warpAffine(frame_bgr, M, (W, H))[:, :, ::-1] /255.0- MEAN) / STD # BGR -> RGB crops += [c, c[:, ::-1]] Ms.append(M) batch = torch.tensor(np.stack(crops).transpose(0, 3, 1, 2).copy(), dtype=torch.float32, device=device)with torch.no_grad(): hm = model(batch).cpu().numpy() out = []for b, M inenumerate(Ms): mirrored = hm[2* b +1][FLIP][:, :, ::-1] # undo mirroring, swap left/right mirrored = np.concatenate([mirrored[:, :, :1], mirrored[:, :, :-1]], axis=2) # 1-pixel shift px, py, conf = peaks((hm[2* b] + mirrored) /2) back = cv2.invertAffineTransform(M) out.append(np.c_[back[0, 0] * px + back[0, 1] * py + back[0, 2], back[1, 0] * px + back[1, 1] * py + back[1, 2], conf])return np.array(out)
Settings (Config)
"""Default settings. Detection/tracking defaults were tuned on hand-checked siamang boxes (11 clips) andDeepWild (Wiltshire et al. 2023; 1577 hand-labelled chimpanzees and bonobos)."""from dataclasses import dataclass, fieldfrom typing import Optional, Tuple# Savitzky-Golay window in seconds (polynomial order 2). "high" lags fast movements such as brachiation.SMOOTHING = {"none": 0.0, "low": 0.1, "medium": 0.25, "high": 0.5}POSE_WEIGHTS_FILE ="openapepose_hrnet_w48_fp16.safetensors"POSE_WEIGHTS_URL ="https://zenodo.org/records/22989835/files/openapepose_hrnet_w48_fp16.safetensors?download=1"POSE_WEIGHTS_SHA256 ="9ac7636c3b086367a9728164537642c5d87020f1337d181e4ba779330afea8c7"@dataclassclass Config:# --- detection: OWLv2 (Minderer et al. 2023), open-vocabulary, prompted with text --- detector_model: str="google/owlv2-base-patch16-ensemble" prompts: Tuple[str, ...] = ("a photo of a chimpanzee", "a photo of a bonobo", "a photo of a gorilla","a photo of an orangutan", "a photo of a gibbon", "a photo of a siamang","a photo of an ape") detection_threshold: float=0.15# keep boxes scoring at least this nms_iou: float=0.7# merge boxes that overlap more than this tiles: int=1# >1: also detect on tiles x tiles enlarged crops (far-away apes; slower) detect_every: int=1# detect on every n-th frame only (CPU); boxes in between are interpolated# --- tracking: ByteTrack (Zhang et al. 2022, via supervision) + stitching + gap filling --- track_activation_threshold: float=0.25# score needed to start a track track_buffer_s: float=1.0# how long a lost track is kept track_matching_threshold: float=0.9# 1 - IoU needed to link a box to a track (lenient: fast apes) stitch_gap_s: float=2.0# re-join a track that reappears within this time (occlusion) stitch_distance: float=1.0# ... within this many ape sizes of where it disappeared fill_gap_s: float=0.5# interpolate missing boxes within a track up to this long# --- smoothing --- box_smoothing: str="low"# none | low | medium | high keypoint_smoothing: str="low"# none | low | medium | high# --- pose: OpenApePose (Desai et al. 2023) --- pad: float=0.1# context around the box given to the pose model pose_weights: Optional[str] =None# local .safetensors file; None = download on first use# --- output video --- draw: Tuple[str, ...] = ("points",) # any of "points", "skeleton"; () = no video output_width: Optional[int] =None# None = same size as the input min_likelihood_draw: float=0.0# hide keypoints/limbs below this likelihood device: Optional[str] =None# None = cuda if available, else cpu extra: dict= field(default_factory=dict)
Tracking: ByteTrack, stitching, gap filling
def bytetrack(dets, fps, threshold, nms_iou, activation, buffer_s, matching): every =int(dets.every.iloc[0]) if"every"in dets andlen(dets) else1 n_frames =int(dets.n_frames.iloc[0]) iflen(dets) else0 rate = fps / every tr = sv.ByteTrack(track_activation_threshold=activation, lost_track_buffer=max(1, int(round(buffer_s * rate))), minimum_matching_threshold=matching, frame_rate=max(1, int(round(rate)))) tr.reset() # track IDs start at 1 in every video by_frame = {f: g for f, g in dets[dets.conf >= threshold].groupby("frame")} rows = []for i inrange(0, n_frames, every): g = by_frame.get(i) d = sv.Detections(xyxy=g[["x0", "y0", "x1", "y1"]].to_numpy(float), confidence=g.conf.to_numpy(float), class_id=np.zeros(len(g), int)) if g isnotNoneelse sv.Detections.empty()iflen(d): d = d.with_nms(threshold=nms_iou, class_agnostic=True) d = tr.update_with_detections(d)for (x0, y0, x1, y1), tid, c inzip(d.xyxy, d.tracker_id, d.confidence): rows.append({"frame": i, "time_s": i / fps, "track_id": int(tid), "det_conf": float(c), "interpolated": False,"box_x": x0, "box_y": y0, "box_w": x1 - x0, "box_h": y1 - y0})return pd.DataFrame(rows, columns=TRACK_COLS)def stitch_tracks(boxes, fps, max_gap_s, max_dist=1.0, max_size_ratio=2.0):"""A track that starts 1 frame..max_gap_s after another ended, within max_dist x ape size (sqrt box area) of where that one ended and at a similar size, gets the older track's ID (an ape that was hidden and reappeared). Joins are one-to-one, closest first, so two apes cannot take over each other's ID."""if max_gap_s <=0ornotlen(boxes):return boxes b = boxes.sort_values("frame") ends = b.groupby("track_id").tail(1).set_index("track_id") starts = b.groupby("track_id").head(1).set_index("track_id") cand = []for old, e in ends.iterrows(): es = np.sqrt(e.box_w * e.box_h)for new, st in starts.iterrows(): gap = st.frame - e.frameif new == old or gap <1or gap > max_gap_s * fps:continue ss = np.sqrt(st.box_w * st.box_h) dist = np.hypot(st.box_x + st.box_w /2- e.box_x - e.box_w /2, st.box_y + st.box_h /2- e.box_y - e.box_h /2)if dist < max_dist * es andmax(ss, es) /min(ss, es) < max_size_ratio: cand.append((dist / es, old, new)) joined, used_old, used_new = {}, set(), set()for _, old, new insorted(cand):if old notin used_old and new notin used_new: joined[new] = old used_old.add(old) used_new.add(new)def root(t):while t in joined: t = joined[t]return treturn boxes.assign(track_id=boxes.track_id.map(root))def fill_gaps(boxes, fps, max_gap_s):"""Linearly interpolate a track's box over gaps up to max_gap_s (rows get interpolated=True).""" max_gap =int(round(max_gap_s * fps))if max_gap <1ornotlen(boxes):return boxes new = []for tid, g in boxes.sort_values("frame").groupby("track_id"): f = g.frame.to_numpy()for a in np.flatnonzero(np.diff(f) >1):if f[a +1] - f[a] -1> max_gap:continue ra, rb = g.iloc[a], g.iloc[a +1]for fr inrange(f[a] +1, f[a +1]): t = (fr - f[a]) / (f[a +1] - f[a]) row = {c: ra[c] + t * (rb[c] - ra[c]) for c in BOX_COLS} row.update(frame=fr, time_s=fr / fps, track_id=tid, det_conf=np.nan, interpolated=True) new.append(row) out = pd.concat([boxes, pd.DataFrame(new, columns=TRACK_COLS)], ignore_index=True) if new else boxesreturn out.sort_values(["frame", "track_id"]).reset_index(drop=True)
Smoothing: boxes and keypoints
"""Savitzky-Golay smoothing (order 2) per track, over runs of consecutive frames only (never across frames wherethe ape was not tracked). Window = SMOOTHING[level] seconds, at least 5 frames (order 2 needs 5 to smooth)."""import numpy as npfrom scipy.signal import savgol_filterfrom .config import SMOOTHINGdef _window(level, fps):if SMOOTHING[level] ==0:return0returnmax(int(round(SMOOTHING[level] * fps)) |1, 5)def _smooth_runs(df, cols, win):"""Smooth columns `cols` of df (sorted by track, frame) within runs of consecutive frames.""" sig = df[cols].to_numpy(float) run = ((df.frame.diff() !=1) | (df.track_id.diff() !=0)).cumsum().to_numpy()for r in np.unique(run): idx = np.flatnonzero(run == r) w =min(win, len(idx) iflen(idx) %2elselen(idx) -1)if w >=5: sig[idx] = savgol_filter(sig[idx], w, 2, axis=0)return sigdef smooth_boxes(boxes, fps, level):"""Smooths each box's centre, width and height. Keeps the unsmoothed box as raw_box_*.""" out = boxes.sort_values(["track_id", "frame"]).reset_index(drop=True)for c in ["box_x", "box_y", "box_w", "box_h"]: out["raw_"+ c] = out[c] win = _window(level, fps)ifnot win ornotlen(out):return out out["cx"], out["cy"] = out.box_x + out.box_w /2, out.box_y + out.box_h /2 s = _smooth_runs(out, ["cx", "cy", "box_w", "box_h"], win) out["box_w"], out["box_h"] = np.maximum(s[:, 2], 1), np.maximum(s[:, 3], 1) out["box_x"], out["box_y"] = s[:, 0] - out.box_w /2, s[:, 1] - out.box_h /2return out.drop(columns=["cx", "cy"])def smooth_keypoints(tracks, keypoints, fps, level):"""Smooths every keypoint's x and y over time (likelihoods are left as they are).""" win = _window(level, fps)ifnot win ornotlen(tracks):return tracks out = tracks.sort_values(["track_id", "frame"]).reset_index(drop=True) cols = [f"{k}_{a}"for k in keypoints for a in"xy"] out[cols] = _smooth_runs(out, cols, win)return out
Limitations
Full evaluation still needs to be performed. Distant apes or overlapping apes are not tracked well. It is limited to apes, monkeys will likely not be tracked well.
Evaluation
Planned (if you want to collaborate on this reach out to w.pouw@tilburguniversity.edu)
OpenApe Pose model
We redistributed the trained OpenApePose model (HRNet-W48, file hrnet_w48_oap_256x192_full.pth) on Zenodo (https://zenodo.org/records/22989835), which was originally deposited on Dryad under public domain (CC0) (https://doi.org/10.5061/dryad.c59zw3rds) and here released under the MIT license in the repository (https://github.com/desai-nisarg/OpenApePose). The only change is a conversion to safetensors [fp16], keeping the backbone and keypoint head weights; the model was not retrained or modified. Please cite the original authors (Desai et al., 2023).
Licence
MIT. Model weights: OWLv2 Apache-2.0; OpenApePose weights CC0 license.
Full documentation (installation, usage, outputs, settings, speed): see the README.