This study demonstrates that large language models (LLMs) can be employed as practical research tools in linguistics, even by researchers without a technical background, provided that experimental pipelines are designed with transparency and reproduci...
This study demonstrates that large language models (LLMs) can be employed as practical research tools in linguistics, even by researchers without a technical background, provided that experimental pipelines are designed with transparency and reproducibility in mind. Using Chinese factivity inference as a case study, the paper presents a step-by-step construction of LLM-based experimental workflows under both cloud-based and local environments. Rather than optimizing model performance per se, the study focuses on observing how LLMs process linguistically complex inference tasks and how their judgments change when external linguistic knowledge is introduced. Three experimental settings are systematically compared: a No-RAG baseline, a Plain -RAG model incorporating minimal retrieval augmentation, and a fully traceable cloud pipeline with execution monitoring. In addition, a local pipeline based on Ollama, DeepSeek-R1, and FAISS is implemented to address concerns of data security, cost, and reproducibility. The results show that even simple retrieval augmentation can substantially reduce hallucination and improve interpretability in factivity inference. By framing LLMs as objects of linguistic analysis rather than black-box optimizers, this study offers a concrete methodological guide for linguists seeking to integrate LLMs into empirical research.