01
The problem
Can V-JEPA representations recognize stages, errors and component states in real industrial procedures? Answering this requires separating visual information from temporal shortcuts and testing on participants absent from training.
02
My role
I prepared datasets and evaluation protocols, cached V-JEPA embeddings and compared linear probes, MLPs, added time features and GRUs.
03
Constraints
- Distinguish content learning from temporal shortcuts.
- Separate participants between training and evaluation.
- Interpret results without confusing memorization and generalization.
04
Architecture & pipeline
- Evaluate V-JEPA 2.1 ViT-L on 84 industrial procedure sequences, with roughly 1.6-second windows.
- Compare linear probes, MLPs, added time features and GRUs on stages, actions, errors and component states.
05
Decisions & iterations
Participant-based splits reduce the risk of overstating generalization from similar sequences.
Caching embeddings separates extraction cost from iterations on downstream models.
A temporal baseline tests whether the position in a procedure already explains part of the targets.
A single-sequence sanity check verifies that the pipeline can learn in a constrained setting; it does not measure generalization.
06
Results & limitations
Temporal models improve some sequence tasks, particularly procedure stages and component states.
On some errors, simple time information can compete with or exceed embeddings. A GRU therefore does not automatically solve anomaly detection.
The sanity check reaches nearly 100% on several targets in one sufficiently trained sequence. This validates the pipeline’s learning capacity, not production performance.
07
Next iterations
Compare video representations with temporal baselines for each error type, using participant-based splits and a stable protocol. Identify gains attributable to visual content before choosing a model for anomaly detection.


