Xuling Zhang's CV

Education

Universiti Sains Malaysia (USM)

Master of Computer Science

Oct 2023 – Mar 2025

GPA: 3.3 / 4.0

Relevant Courses:

  • Deep Learning
  • Multimodal Learning
  • Internet of Things
  • Algorithm Design

Yingkou Institute of Technology

Bachelor of Intelligent Science and Technology

Sep 2018 – Jun 2022

GPA: 3.2 / 4.0

Honors:

  • Outstanding Graduate

Relevant Courses:

  • Machine Learning
  • Data Mining
  • Big Data Analytics
  • Data Visualization

Work Experience

360 Digital Security Group

Large Model Algorithm Engineer

Aug 2025 – Present

WatchLog (EDR Log Security Analysis Framework) ICML 2026 Regular

  • Proposed and developed WatchLog, a multimodal LLM framework that transforms ultra-long EDR log sequences (up to 1M tokens) into video-like structures for efficient and interpretable event reasoning.
  • Designed a novel kvEmbedding layer to encode key-value log events into image tensors, and applied temporal cross-attention to compress long sequences into fixed-length representations, significantly reducing computational overhead.
  • Implemented a three-stage training pipeline (image-event alignment, video-log alignment, and supervised fine-tuning) to simultaneously output threat categories and natural language reasoning traces.
  • Achieved 99.8% binary classification accuracy and 100% recall on the custom EDR8M-20R dataset, outperforming traditional methods and fine-tuned LLMs.
  • Reduced inference latency by over 200x (first-token latency of 1.11s) and GPU memory usage by 50.6% (48.6GB for 1024K tokens) compared to Qwen2.5-7B, enabling single-GPU processing of million-token logs.

PSIFrame (Self-Evolving LLM Multi-Agent Framework) NeurIPS 2026 Under Review

  • Developed PSIFrame, a self-evolving multi-agent framework designed to overcome the limitations of static workflows and flat tool selection in dynamic, open-world environments.
  • Constructed a hierarchical Taxonomy Graph for tool organization, enabling top-down, task-aware routing through functional and executable agents.
  • Integrated Monte Carlo Tree Search (MCTS) with an adversarial debate module for adaptive planning, execution, and dynamic backtracking based on environmental feedback.
  • Improved tool retrieval NDCG@10 by 21.3% and Recall@10 by 19.8% on the MetaTool dataset, while reducing token consumption by ~50%.
  • Achieved an 80.0% task success rate on a custom open-world multi-agent benchmark (BenchVerse), significantly outperforming baselines like MetaAgent (68.0%).

NDR Large Model Distillation (Ongoing)

  • Conducting data analysis on NDR log JSON files and fine-tuning Qwen3-14B using LlamaFactory with a CoE architecture.
  • Optimizing algorithms to address long-tail distribution problems, successfully improving model performance from 82% to 93% to date.

SF Technology

AI Engineer Intern

Feb 2025 – Jun 2025

Project

Multimodal Financial Shared-Service LLM System

Technologies

  • Qwen
  • LoRA
  • SFT
  • OCR
  • YOLOv5
  • PyTorch

Contributions

  • Fine-tuned multimodal LLMs (e.g., Qwen-VL, LLaVA) using LoRA and SFT for logistics and financial document understanding, improving information extraction accuracy from 85% to 96%.
  • Built and optimized OCR pipelines using YOLOv5 and CRNN for complex invoice and receipt parsing, increasing parsing precision by 15% and reducing processing time per document by 40%.
  • Applied object detection models for automated cargo verification and anomaly detection, enhancing the robust performance of the vision system under varying lighting conditions.
  • Integrated multimodal reasoning capabilities into financial shared-service workflows, reducing manual verification workload by 60%.

Lazada

Data Mining & Recommendation Intern

Jun 2024 – Dec 2024

Technologies

  • PCA
  • LDA
  • SVM
  • KNN
  • XGBoost
  • Collaborative Filtering

Contributions

  • Conducted deep-dive CTR and CVR analysis on user behavior data to identify key features driving user engagement and conversion.
  • Built and deployed robust recommendation and ranking models using XGBoost and LightGBM, improving overall CTR by 12% and CVR by 8%.
  • Implemented collaborative filtering algorithms (UserCF, ItemCF) and matrix factorization techniques (SVD, NMF) to alleviate the cold-start problem, increasing recommendation coverage by 20%.
  • Applied Learning-to-Rank (L2R) techniques to optimize the ranking phase of the recommendation pipeline, yielding a 5% improvement in NDCG@10.

Research Experience

Hong Kong University of Science and Technology (Guangzhou)

Research Assistant

Jan 2025 – Aug 2025

Topic

Replay-Free Continual Graph Learning

Research Problem

Graph neural networks suffer from catastrophic forgetting when new classes or tasks arrive continuously.

Contributions

  • Designed an analytic learning framework for continual graph learning.
  • Eliminated replay buffers to preserve privacy and reduce storage cost.
  • Developed recursive least squares based memory updates.
  • Constructed recursive correlation memory matrices for stable knowledge retention.
  • Improved training efficiency while maintaining competitive performance.

Representative Work

AL-GNN: Privacy-Preserving and Replay-Free Continual Graph Learning via Analytic Learning


Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences

Research Collaborator

Dec 2024 – Jun 2025

Topic

Multimodal Large Language Model Alignment

Research Problem

Existing multimodal foundation models show weak generalization on scientific and chemical reasoning tasks.

Contributions

  • Built a benchmark covering eight multimodal chemistry tasks.
  • Investigated DPO and GRPO based multimodal alignment strategies.
  • Improved cross-modal consistency through balanced optimization.
  • Enhanced transferability across heterogeneous reasoning tasks.

Technical Skills

Programming Languages

Frameworks

Research Areas

Tools


Awards and Honors

Research

Competitions

Scholarships

Honors


Academic Vision

My long-term research vision centers on Multimodal World Models and Continual AI Systems. I believe the next frontier of artificial intelligence lies at the intersection of multimodal perception and persistent learning — building agents that can continuously perceive, reason, and adapt across visual, linguistic, and structured modalities without forgetting previously acquired knowledge.

Specifically, I am passionate about:

I envision a future where AI systems can learn continuously from multimodal streams of experience, build coherent world models, and generalize across tasks and domains — much like humans do. This pursuit drives my work at the intersection of graph learning, multimodal reasoning, and continual adaptation.


Publications

======