AI Engineer & Researcher · Shanghai

Portrait of Rui Wang

Rui Wang

I build large language model systems that hold up in production, and measure exactly where they fail.

15+ years shipping search, ranking, and LLM systems. Technical Manager at Boai (China) Enterprise Group. M.Eng., Arizona State University.

fig. 1 level-wise BDER (%) · zero-shot
Depth Qwen3 Qwen3.5 Llama-3.1 InternLM3† NB
Chapterl=3 58.964.775.999.9 7.2K
Sectionl=4 30.742.647.498.1 0.9K
Articlel=5 8.711.38.4100.0 52.5K
Clausel=6 23.022.025.497.0 182.5K
Iteml=7 6.26.67.418.6 1,085.6K
Sub-iteml=8 7.67.69.196.8 572.1K

58.9% of Chapter boundaries are missed even by the best base model.

Table IV of the StructEval paper. Boundary detection error rate by hierarchy depth on a 130.2M-token Chinese legal corpus, 8–9B models; lower is better. † instruction-tuned. N_B: number of boundary tokens.
15+yrs
engineering in production
10B+
user behavior logs in production
14B+
parameter medical LLM, open weights
IEEE
first-author paper, ITASC 2026

§01

Research

Long-context benchmarks ask whether a model can find a fact. I study whether it understands how a document is built.

Conference paper IEEE ITASC 2026 Fukuoka, Japan pp. 52–57

StructEval: A Systematic Evaluation Framework for Structural Awareness in Large Language Models

Rui Wang, Yinong Chen, Hengkuan Fu, Yuemin Bian

Arizona State University · Shanghai Yida Hospital · Digital Medical Research Institute, Shanghai University

LLMs handle long contexts well but often lose track of documents with deep, nested hierarchy. StructEval uses a deterministic parser to extract boundaries, depths, and types from raw text, giving token-level ground truth for structural prediction. Tested on a Chinese legal corpus of over 130 million tokens.

  1. F1 Structure blindness is universal. Every model tested (Llama-3.1, Qwen-3/3.5, InternLM-3) misses structure zero-shot. Base models are weakest at high-level boundaries.
  2. F2 Alignment costs topology. Instruction-tuned models lose structural awareness, and the loss grows with context length.
  3. F3 Reproducible by design. Labels come from a deterministic pipeline, not from human or model judgment.

§02

Yida-Model-14B

Not a bigger model. A model that knows medicine. An open-weight Chinese medical LLM built on Qwen3-14B, where most of the work went into the data.

1,000+ medical textbooks and clinical reference books in the source corpus
350K curated fine-tuning samples after quality filtering
3 scoring axes: medical accuracy, expression, safety & ethics
#10 tied, MedBench public leaderboard (research snapshot)
  1. 01IngestPDF / OCR, structure recovery
  2. 02Synthesizemulti-model QA generation
  3. 03Governscoring, balancing, vector dedup
  4. 04TuneLoRA r16 on 4 GPUs, merged to bf16
  5. 05EvaluateMedBench-driven gap filling

A research model for evaluation. Not a medical device, and not a basis for clinical decisions.

§03

Experience

From game back ends to legal search to medical LLMs. The constant is systems that real users depend on.

  1. Jan 2024 — Present

    Technical Manager current

    Boai (China) Enterprise Group

    Shanghai University AI Medical Joint Innovation R&D Center, a research center run with Shanghai University

    • Led Yida-Model-14B, a Chinese medical LLM: LoRA-based SFT on an open-source base model, from corpus construction to merged, published weights.
    • Owned training data cleaning: OCR ingestion, multi‑model QA synthesis, three‑axis scoring, and semantic deduplication.
    • Built RAG question answering and local knowledge bases on the fine-tuned models, with agent tool use and function calling.
    • Deployed, evaluated, and tested open‑source LLMs on private infrastructure, and wrote the evaluation reports.
    • Researched and deployed distributed training and inference services for large models.
    • Led model validation for the team, trained other engineers, and brought new AI research into production.
  2. Jun 2017 — Dec 2023

    Senior Search Algorithm Engineer

    LexisNexis

    • Developed search engine features for legal research products.
    • Trained learning-to-rank models (RankLib, LambdaMART) to rerank search results, validated with A/B tests.
    • Built legal NER for case numbers, document numbers, penalty numbers, and company names, and benchmarked it against LLM extraction.
    • Built document recommendation for related statutes, interpretations, and similar cases, trained on user behavior logs.
    • Built Hyperlink, which analyzes and links citations across billions of documents, updated daily.
    • Built user behavior analytics over tens of billions of log records on Elasticsearch, ClickHouse, and Spark.
    • Tuned Solr search clusters and Spark clusters; built tokenizers, scoring systems, and data visualization.
  3. Apr 2016 — May 2017

    Senior Big Data Engineer

    Wealink

    • Processed and analyzed resume data with NLP.
    • Developed resume search business logic on Elasticsearch.
    • Ran Spark analytics, reporting, visualization, and big data storage.
    • Debugged data crawlers and designed high‑concurrency web architecture.
  4. May 2015 — Apr 2016

    Technology Partner (Co-founder)

    Hanyou Network

    • Co-founded a mobile game platform with a small group of partners.
    • Ran the engineering department and owned the product technical architecture.
    • Managed the technical team and solved its hardest technical problems.
  5. Mar 2012 — May 2015

    Senior Python Engineer

    Kingnet

    • Developed systems for domestic and international games.
    • Debugged and integrated game back-end APIs.
    • Developed game website front ends.
    • Built the game log collection system and analyzed its data.

§04

Range of work

Nine areas where I have built and shipped systems.

  1. 01

    Large Language ModelsLLM

    Fine-tuned, merged, and published domain large language models, including a Chinese medical model.

    • LLM
    • SFT
    • Qwen
  2. 02

    Machine LearningML

    Trained ranking, classification, and prediction models, including learning-to-rank for search.

    • LTR
    • Ranking
    • Classification
  3. 03

    AI Agent

    Built agents that call tools and complete tasks inside existing business systems.

    • Agent
    • Tool Use
    • Workflow
  4. 04

    Natural Language ProcessingNLP

    Extracted, summarized, and generated text from records, reports, and tables.

    • NLP
    • Summarization
    • OCR
  5. 05

    Information RetrievalIR

    Analyzed citation relationships across billions of documents and linked related articles.

    • Search
    • Citation
    • Ranking
  6. 06

    Computer VisionCV

    Analyzed images and video for counting, monitoring, and visual search.

    • Computer Vision
  7. 07

    Data Engineering

    Built data collection, governance, and business data platforms across operational systems.

    • Data Platform
    • ETL
    • Governance
  8. 08

    Retrieval-Augmented GenerationRAG

    Built retrieval-augmented generation for question answering over private knowledge bases.

    • RAG
    • Knowledge Base
    • Retrieval
  9. 09

    AIOps

    Ran monitoring and operations for services, with metrics, logs, and dashboards.

    • AIOps
    • Grafana
    • Elasticsearch

§05

Working on LLM training, medical AI, or search? Let's talk.

Research collaboration, engineering roles, and technical discussions are all welcome.

contact@wangrui.io