WatchLog: From a Glimpse to Decision—Rapid Event Reasoning in Endpoint Detection and Response Logs with Multimodal LLMs
Overview
Problem & Motivation
EDR systems generate massive logs. Existing methods struggle with limited context windows, lack of interpretability, and temporal sparsity of malicious activities.
Our Approach
We propose WatchLog, a multimodal LLM framework that converts logs into video-like representations. It uses a three-stage pipeline: image–event and video–log pre-training, followed by supervised fine-tuning, enabling interpretable reasoning over million-token logs.
Results
WatchLog improves detection BA from 93.3% (Qwen3-235B) to 99.8%, and OA from 51.3% (Qwen2.5-7B) to 90.3%, while maintaining low GPU memory and latency.

Authors and Affiliations
- Hongyi Zhou — Tsinghua University, China
- Jianfeng Pan — 360 Digital Security Group, China
- Min Peng — 360 Digital Security Group, China
- Shaomang Huang — China Telecom Cloud Technology, China
- Xuling Zhang — 360 Digital Security Group, China
Research Area
- AI Security
- Endpoint Detection and Response
- Multimodal LLMs
- Explainable Threat Detection
Key Contributions
- Converted ultra-long EDR logs into video-like multimodal representations for efficient reasoning.
- Designed a three-stage training paradigm combining image-event pre-training, video-log alignment, and supervised fine-tuning.
- Improved model interpretability by producing understandable reasoning traces over long event sequences.
- Achieved strong efficiency-performance trade-offs for enterprise-grade endpoint security analysis.
Publication Details
- Venue: International Conference on Machine Learning 2026 (ICML 2026)
- Presentation Type: Regular
Download
Paper link will be added later.
