跳过正文
  1. 政治学英文顶刊消息速递/
  2. 论文时间线/

Evaluating the Effectiveness of Open‐Source LLMs for Automated Analysis of Multilingual Consultation Feedback: A Swiss Case Study

PSJ 54/4

2026-09-14

Original

Evaluating the Effectiveness of Open‐Source LLMs for Automated Analysis of Multilingual Consultation Feedback: A Swiss Case Study

Edgar Mathevet, Lauriane Cailleux

Abstract

ABSTRACT Public consultations are a cornerstone of democratic policymaking, yet the sheer volume and heterogeneity of submissions often overwhelm traditional analysis. Emerging large language models (LLMs) offer new opportunities for making sense of such data, but their performance in contexts with high linguistic and technical diversity remains underexplored. In this research note, we present a case‐based demonstration rather than a generalizable evaluation, assessing an open‐source LLM (Gemma3:27b) and a proprietary LLM (GPT‐5‐mini) for analyzing stakeholder inputs from three Swiss pre‐parliamentary consultations. The coexistence of multiple languages and varying technical complexity in submissions creates uniquely challenging datasets, providing a rigorous test for automated tools. We benchmark the model against official administrative reports, assessing accuracy and completeness. Results show the open‐source model performs excellently in straightforward consultations, achieving a mean accuracy of 90%, but encounters difficulties in retrieving the substantive content (overall completeness: 70%). Comparatively, the proprietary model obtains better results than the open‐source one in these two settings. Overall, LLMs can streamline analysis and enhance interpretive capacity, complementing (but not replacing) human expertise. The study highlights both the potential and limits of algorithmic support for participatory governance in linguistically and technically diverse environments.

中文

评估开源大语言模型在多语言咨询反馈自动分析中的有效性:基于瑞士的案例研究

Edgar Mathevet, Lauriane Cailleux

摘要

公众咨询是民主政策制定的基石,但提交材料在数量和异质性上的庞大规模常常使传统分析不堪重负。新兴的大语言模型(LLMs)为理解此类数据提供了新机遇,但其在语言和技术高度多样化情境中的表现仍未被充分探索。在本研究笔记中,我们呈现的是一项基于案例的演示,而非可推广的评估,旨在评估一个开源大语言模型(Gemma3:27b)和一个专有大语言模型(GPT‐5‐mini)对瑞士三项议会前咨询中利益相关者意见的分析能力。提交材料中多种语言并存以及技术复杂度不一,构成了极具挑战性的数据集,为自动化工具提供了严格测试。我们以官方行政报告为基准,评估模型的准确性和完整性。结果显示,开源模型在较为直接的咨询中表现优异,平均准确率达到90%,但在提取实质性内容方面遇到困难(总体完整性为70%)。相比之下,专有模型在这两种情境中均取得优于开源模型的结果。总体而言,大语言模型能够简化分析并增强解释能力,可补充但不能取代人类专业知识。该研究凸显了算法支持在语言和技术多样化环境中参与式治理的潜力与局限。

关键词

大语言模型、公众咨询、多语言分析、自动化文本分析、参与式治理、瑞士、政策制定、模型评估