(en) Disentangled Representation Learning seeks to recover the latent factors of variation that generate high-dimensional observations, providing a foundation for interpretable, robust machine learning. Despite empirical progress, the field is marked by a gap between heuristic architectures and the formal requirements of identifiability theory. This thesis bridges that gap with a unified framework for principled disentanglement, grounded in functional independence, geometric constraints on the generative mapping, and the structure of sequential data. Its central claim is that what makes a latent variable a meaningful factor is a geometric property of the generative mapping rather than a statistical property of the data. This yields formal identifiability guarantees that hold without assuming independence between factors, and explains why certain model families succeed at disentanglement while others fail. Sequential data plays a privileged role: the temporal structure of video offers a natural source of interventional information. The thesis shows that static and dynamic factors are generically causally dependent, and that respecting this asymmetry is the key to provable disentanglement without supervision. The same principles further yield compact global representations of video sequences, uniting identifiability and informativeness under a single design principle.