Abstract
Animal re-identification is the process of recognising individual animals across different images or video frames, often captured at varying times or locations. Unlike general object detection, which identifies the presence of an animal, re-identification focuses on distinguishing one specific animal from others of the same species. This task is important in ecological monitoring, wildlife conservation, and behavioural studies, where tracking individuals over time provides insights into movement patterns, social interactions, and health status. Due to differences in appearance, pose, lighting, and occlusion, animal re-identification is a challenging problem that often requires specialised datasets and tailored machine learning techniques.The emergence of machine learning and computer vision has facilitated the automation of animal re-identification, predominantly through supervised learning frameworks and deep learning architectures that depend on manually annotated datasets to achieve consistent and reliable performance. However, the availability of such datasets remains limited, highlighting the necessity for additional benchmark datasets to support the development and evaluation of novel animal re-identification methodologies. In response to this need, a multi-species animal video dataset was constructed, incorporating bounding boxes, identity labels, and multiple feature representations. This dataset served both to evaluate the proposed methods in this work and to address the scarcity of benchmark resources within the field.
As animal re-identification solutions are typically designed for a bespoke subpopulation of a single species, there is a clear need for the development of generalisable methodologies. In response, this work presents the design of a fully autonomous species-invariant animal re-identification pipeline capable of operating in both online and offline scenarios. As a foundational step, an experimental study was conducted to identify the most reliable feature representation for the benchmark dataset in the context of animal re-identification. The results indicated that simple RGB-based features were effective for animal re-identification across species, and were consequently employed in all subsequent experiments.
The development of a novel object detection paradigm is introduced, combining outputs from object detection and multiple object tracking techniques via intersection over union thresholding and connected component extraction to enhance detection accuracy and reliability. To assess and manage the structural complexity of the dataset, both hierarchical and centroid-based constrained clustering approaches were evaluated, with hierarchical clustering proving more successful due to the presence of elongated cluster formations commonly observed in the data.
Building on these findings, the work presented in this thesis contributed to the conception and development of both offline and online semi-supervised clustering methods. Two offline approaches are proposed. The first employs hierarchical clustering of object tracks derived from object detection and multi-object tracking algorithms, with tracks subsequently merged based on classification outputs and a resubstitution confusion matrix. The second approach uses a semi-supervised clustering ensemble, in which a set of constrained clustering algorithms is applied to generate a library of base partitions, which are then integrated using a cumulative adjacency matrix. In addition, an online method was designed to support real-time video analysis by incorporating spatio-temporal constraints into an incremental clustering framework and a likelihood thresholding mechanism to distinguish between new and existing identities. This enables streamlined updates and summarisation of clusters while maintaining low memory requirements and high re-identification accuracy. All proposed semi-supervised clustering methods were evaluated against state-of-the-art baselines and consistently demonstrated superior performance in achieving species-invariant animal re-identification across the benchmark video dataset.
| Date of Award | 23 Sept 2025 |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Sponsors | Artificial Intelligence, Machine Learning and Advanced Computing (AIMLAC) CDT |
| Supervisor | Ludmila Kuncheva (Supervisor) |
Keywords
- Animal Re-Identification
- Constrained Clustering
- Re-Identification Pipeline
- Online Clustering
- Unrestricted Videos
- Benchmark Dataset
- PhD
Cite this
- Standard