StructEval: A Systematic Evaluation Framework for Structural Awareness in Large Language Models
Arizona State University · Shanghai Yida Hospital · Digital Medical Research Institute, Shanghai University
LLMs handle long contexts well but often lose track of documents with deep, nested hierarchy. StructEval uses a deterministic parser to extract boundaries, depths, and types from raw text, giving token-level ground truth for structural prediction. Tested on a Chinese legal corpus of over 130 million tokens.
- F1 Structure blindness is universal. Every model tested (Llama-3.1, Qwen-3/3.5, InternLM-3) misses structure zero-shot. Base models are weakest at high-level boundaries.
- F2 Alignment costs topology. Instruction-tuned models lose structural awareness, and the loss grows with context length.
- F3 Reproducible by design. Labels come from a deterministic pipeline, not from human or model judgment.
@inproceedings{wang2026structeval,
author = {Wang, Rui and Chen, Yinong and Fu, Hengkuan and Bian, Yuemin},
title = {{StructEval}: A Systematic Evaluation Framework for
Structural Awareness in Large Language Models},
booktitle = {2026 IEEE International Symposium on Intelligent
Transportation and Smart City (ITASC)},
address = {Fukuoka, Japan},
pages = {52--57},
year = {2026},
doi = {10.1109/ITASC70892.2026.00018}
}