MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education | TickerVault