TY - GEN
T1 - Prompt Engineering Based Factivity Hallucination Alleviation System for Large Language Models
AU - Yang, Heng
AU - Zeng, Na
AU - Cui, Tianxiang
AU - Lin, Qiao
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026/7
Y1 - 2026/7
N2 - Large Language Models (LLMs) show excellent abilities in various applications. However, they tend to produce hallucinations - -fabricated or factually incorrect information, which reduces their reliability. Therefore, LLMs hallucination has become an increasingly prominent problem, urgently requiring improvements in the factual consistency of model outputs. This paper proposes a prompt engineering framework that can comprehensively alleviate the fact hallucination. The proposed approach combines dynamic threshold mechanisms, cross-validation, self-verification, chain-of-thought reasoning, and retrieval-augmented generation techniques. Extensive experiments were conducted using the Claude 3.7 and Deepseek-V3 models across various domains, including medical, arts, entertainment, and others. Experimental results show that the proposed prompt framework achieves 2.23% to 4.94% accuracy improvements over baseline models without prompt optimization, while F1 scores improve by approximately 7.09%. The framework shows particular effectiveness in reducing factual hallucinations while maintaining model performance across various query types. Our findings establish a sustainable framework for research addressing hallucinations in LLMs through prompt engineering, while also delivering practical insights for deploying more reliable AI systems.
AB - Large Language Models (LLMs) show excellent abilities in various applications. However, they tend to produce hallucinations - -fabricated or factually incorrect information, which reduces their reliability. Therefore, LLMs hallucination has become an increasingly prominent problem, urgently requiring improvements in the factual consistency of model outputs. This paper proposes a prompt engineering framework that can comprehensively alleviate the fact hallucination. The proposed approach combines dynamic threshold mechanisms, cross-validation, self-verification, chain-of-thought reasoning, and retrieval-augmented generation techniques. Extensive experiments were conducted using the Claude 3.7 and Deepseek-V3 models across various domains, including medical, arts, entertainment, and others. Experimental results show that the proposed prompt framework achieves 2.23% to 4.94% accuracy improvements over baseline models without prompt optimization, while F1 scores improve by approximately 7.09%. The framework shows particular effectiveness in reducing factual hallucinations while maintaining model performance across various query types. Our findings establish a sustainable framework for research addressing hallucinations in LLMs through prompt engineering, while also delivering practical insights for deploying more reliable AI systems.
KW - hallucination
KW - large language model
KW - prompt engineering
UR - https://www.scopus.com/pages/publications/105044862316
U2 - 10.1109/CSIS-IAC70275.2026.11585163
DO - 10.1109/CSIS-IAC70275.2026.11585163
M3 - Conference contribution
AN - SCOPUS:105044862316
T3 - 2026 International Annual Conference on Complex Systems and Intelligent Science, CSIS-IAC 2026
SP - 27
EP - 32
BT - 2026 International Annual Conference on Complex Systems and Intelligent Science, CSIS-IAC 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2026 International Annual Conference on Complex Systems and Intelligent Science, CSIS-IAC 2026
Y2 - 15 May 2026 through 17 May 2026
ER -