— p. 1 — 2025 UAP Workshop:
Narrative Data, Infrastructures, and Analysis
Workshop Synthesis and Recommendations
August 5-6, 2025
Associated Universities, Inc. (AUI)
Workshop sponsored by:
All-domain Anomaly Resolution Office (AARO)26-P-0344
— 第 1 页 — 2025 UAP 研讨会:
叙述性数据、基础设施与分析
研讨会综述与建议
2025 年 8 月 5-6 日
联合大学公司(Associated Universities, Inc., AUI)
研讨会主办方:
全域异常解决办公室(All-domain Anomaly Resolution Office, AARO)
26-P-0344
— p. 2 — Table of Contents
Executive Summary ........................................................................................................................ 2
Introduction and Purpose ................................................................................................................ 3
About the Workshop ....................................................................................................................... 3
Establishing open dialogue ......................................................................................................... 4
Workshop Summary ....................................................................................................................... 4
Agenda overview ........................................................................................................................ 4
Breakout Discussion Summaries ................................................................................................ 5
Breakout Session #1: Identifying, accessing, and integrating data sources [DAY 1] ............ 5
Breakout Session #2: Pathways for data analysis and interpretation at scale [DAY 1] ......... 5
Breakout Session #3: Cleaning, organizing, and linking data: What can and should be done?
[DAY 2] .................................................................................................................................. 5
Outcomes and Recommendations ................................................................................................... 7
Synthesis of Findings .................................................................................................................. 7
Relevant data types and sources of UAP narrative reports ..................................................... 7
Barriers and challenges in data collection and use ................................................................. 7
Metadata and context for usability and analysis ..................................................................... 8
Linking data sources and developing a unified approach ....................................................... 8
Assessing credibility and quality of reports ............................................................................ 8
AI and analytical methods ...................................................................................................... 9
Forward-looking strategy and key considerations .................................................................. 9
Recommended actionable next steps ........................................................................................ 10
Appendix A: Invitation Letter ....................................................................................................... 11
Appendix B: Guidelines for Conduct ........................................................................................... 13
Appendix C: Workshop Agenda ................................................................................................... 14
Appendix D: Breakout Session Prompts....................................................................................... 15
— 第 2 页 — 目录
执行摘要 .......... 2
引言与目的 .......... 3
关于本次研讨会 .......... 3
建立开放对话 .......... 4
研讨会综述 .......... 4
议程概览 .......... 4
分组讨论综述 .......... 5
分组讨论 #1:识别、访问和整合数据来源【第 1 天】 .......... 5
分组讨论 #2:大规模数据分析与解读的路径【第 1 天】 .......... 5
分组讨论 #3:清理、组织与关联数据:可以做什么、应该做什么?【第 2 天】 .......... 5
成果与建议 .......... 7
发现的综合 .......... 7
UAP 叙述性报告的相关数据类型与来源 .......... 7
数据收集与使用中的障碍与挑战 .......... 7
用于可用性与分析的元数据与背景信息 .......... 8
关联数据来源并发展统一方法 .......... 8
评估报告的可信度与质量 .......... 8
AI 与分析方法 .......... 9
前瞻性战略与关键考量 .......... 9
建议的可执行后续步骤 .......... 10
附录 A:邀请函 .......... 11
附录 B:行为准则 .......... 13
附录 C:研讨会议程 .......... 14
附录 D:分组讨论提示 .......... 15
— p. 3 — 2
Executive Summary
From both government and scientific perspectives, advancing Unidentified Anomalous
Phenomena (UAP) research requires rigorous data collection, standardization, and analysis.
Most UAP reports are fragmented, sparse, and unstructured, ranging from military logs and pilot
reports to archival records, social media posts, and civilian testimony. Interpreting this
heterogeneous data at scale is complicated by barriers of classification, translation, and retention.
At the same time, UAP reports also present opportunities for novel methods of integration,
metadata design, and analysis. The 2025 UAP Workshop on Narrative Data, Infrastructures, and
Analysis brought together 40 participants from government, academia, and independent research
organizations. The meeting focused specifically on the challenges and opportunities of working
with UAP narrative reports and related data sources.
Workshop discussions highlighted several cross-cutting findings. First, effective progress
requires clear standards and common reporting templates, with robust metadata capturing time,
location, provenance, morphology, and contextual details. Second, linking across datasets –
military and civilian, to include archival, environmental, and technical - must balance
interoperability with privacy, ethical, and classification constraints. Third, credibility is best
assessed through corroboration, but for efficiency there is a need for automated methods to filter
reports and surface the most promising for investigation. Fourth, AI and machine learning tools
offer capacity for transcription, triage, clustering, and semantic search, but they must be
deployed cautiously to avoid hallucination, bias, and amplification of hoaxes. Human oversight
and iterative workflows remain essential. Finally, the workshop underscored the importance of
community engagement and trust-building, encouraging the scientific community to cultivate a
sustainable “community of practice” for UAP research with further work and convenings.
This report concludes with recommended actionable next steps to establish metadata templates;
combine human expertise with AI tools; leverage existing tools and infrastructures; support
triage with awareness of bias; convene community members; facilitate qualitative integration in
investigation, such as interviews; prioritize collection of new high-quality reports while
integrating historical data; and improve reporting interfaces to enhance accessibility,
collaboration, and transparency. Together, these findings and recommendations point toward a
multi-disciplinary and community-engaged approach to UAP narrative data, which may
influence how and where technical sensors are deployed.
— 第 3 页 — 执行摘要
无论从政府还是科学的视角,推进不明异常现象(UAP)研究都需要严格的数据收集、标准化与分析。大多数 UAP 报告是碎片化、稀疏且非结构化的,其范围从军事日志和飞行员报告,到档案记录、社交媒体帖子和民间证词。对这类异质数据进行大规模解读,因保密、翻译和留存等障碍而变得复杂。与此同时,UAP 报告也为整合、元数据设计与分析的新方法提供了机会。这场关于叙述性数据、基础设施与分析的 2025 UAP 研讨会汇集了来自政府、学术界和独立研究机构的 40 名参与者。会议特别聚焦于处理 UAP 叙述性报告及相关数据来源的挑战与机遇。
研讨会讨论凸显了若干跨领域的发现。第一,要取得有效进展,需要明确的标准和通用的报告模板,并配以能够捕捉时间、地点、来源、形态和背景细节的稳健元数据。第二,跨数据集的关联——军事与民用,包括档案、环境与技术数据——必须在互操作性与隐私、伦理及保密约束之间取得平衡。第三,可信度最好通过相互印证来评估,但为提高效率,需要自动化方法来筛选报告并凸显最有调查价值的报告。第四,AI 和机器学习工具在转录、分流、聚类和语义检索方面提供了能力,但必须谨慎部署,以避免幻觉、偏见和对骗局的放大。人工监督与迭代工作流程仍然必不可少。最后,研讨会强调了社群参与和信任建设的重要性,鼓励科学界通过进一步的工作与聚会,培育一个可持续的 UAP 研究“实践社群”(community of practice)。
本报告最后提出了建议的可执行后续步骤:建立元数据模板;将人类专业知识与 AI 工具相结合;利用现有工具与基础设施;在开展分流时注意偏见;召集社群成员;促进调查中的定性整合(如访谈);在整合历史数据的同时,优先收集新的高质量报告;以及改进报告界面,以增强可及性、协作与透明度。这些发现与建议共同指向一种对 UAP 叙述性数据采取的多学科、社群参与式的方法,这可能会影响技术传感器如何以及在何处部署。
— p. 4 — 3
Introduction and Purpose
Understanding the nature of Unidentified Anomalous Phenomena (UAP) has emerged in recent
years as a pressing area of inquiry in need of rigorous scientific approaches, as well as cross-
disciplinary, cross-sector and international collaboration. Analyzing reports of UAP related
sightings and experiences presents unique challenges due to the large-scale, heterogeneous, and
qualitative nature of the reports originating from military and civilian sources. These reports
typically lack standardized metadata, making comparative analysis difficult. Additionally, the
integration of UAP reports from disparate sources—such as military databases, online reporting
systems, digital and digitized archival records, and social media—poses significant challenges
for harmonization and verification of data and construction of evidence. The complexity of these
datasets requires innovative data infrastructure solutions to enhance reliability, accessibility, and
interoperability. The workshop explored these challenges and sought strategies to improve UAP
data standardization, integration, and analytical approaches.
Recent advances in artificial intelligence (AI) and machine learning present both opportunities to
address challenges, along with potential hazards. Tools such as Large Language Models (LLMs)
can assist with transcription, clustering, and pattern detection at scale, but they risk introducing
bias and hallucination. Responsible use of AI to help organize, analyze, and integrate UAP
reports at scale requires evaluation, human oversight, and shared frameworks for interpretation,
alongside new models to ensure transparency and trust across diverse research communities.
Therefore, the overall purpose of the workshop was to gather perspectives from the broader
scientific community and advance the science of UAP.
About the Workshop
The workshop centered on the collection, organization, and interpretation of UAP reports, with
attention to the challenges and opportunities of working with narrative data. The primary
objectives established for the workshop were to:
• Assess the current landscape of UAP reporting systems and data repositories;
• Identify key challenges and gaps in UAP data collection, standardization, and
accessibility;
• Explore methodologies for data analysis and pattern recognition in UAP reports;
• Nurture trust and collaboration between researchers, government agencies, and civilian
organizations; and
• Propose recommendations for developing a robust UAP data infrastructure.
— 第 4 页 — 引言与目的
近年来,理解不明异常现象(UAP)的本质已成为一个亟需严谨科学方法,以及跨学科、跨部门和国际协作的紧迫探究领域。分析与 UAP 相关的目击和体验报告面临独特的挑战,因为这些源自军方和民间的报告具有大规模、异质和定性的特点。这些报告通常缺乏标准化的元数据,使得比较分析变得困难。此外,整合来自不同来源——如军事数据库、在线报告系统、数字及数字化档案记录、社交媒体——的 UAP 报告,为数据的协调、验证和证据构建带来了重大挑战。这些数据集的复杂性要求创新的数据基础设施解决方案,以增强可靠性、可及性和互操作性。研讨会探讨了这些挑战,并寻求改进 UAP 数据标准化、整合与分析方法的策略。
人工智能(AI)与机器学习的最新进展,既带来了应对挑战的机遇,也带来了潜在的风险。大型语言模型(LLM)等工具可以协助大规模的转录、聚类和模式检测,但存在引入偏见和幻觉的风险。以负责任的方式使用 AI 来大规模地组织、分析和整合 UAP 报告,需要评估、人工监督和共享的解读框架,同时需要新的模式来确保在不同研究社群之间的透明与信任。因此,研讨会的总体目的是汇集更广泛科学界的观点,并推进 UAP 科学。
关于本次研讨会
研讨会围绕 UAP 报告的收集、组织与解读展开,重点关注处理叙述性数据的挑战与机遇。为研讨会确立的主要目标是:
• 评估 UAP 报告系统与数据存储库的现状;
• 识别 UAP 数据收集、标准化与可及性方面的关键挑战与空白;
• 探索 UAP 报告中数据分析与模式识别的方法论;
• 培育研究人员、政府机构与民间组织之间的信任与协作;以及
• 为发展一个稳健的 UAP 数据基础设施提出建议。
— p. 5 — 4
Outside participation was limited due to budget constraints and institutional capacity. Potential
participants were identified based on demonstrated expertise in one or more of the following
areas: AI and machine learning; UAP research and data; physical and natural sciences;
information and data science; archives and records; analysis methods; cyberinfrastructure and
computation; and human and social sciences.
If an invitee declined to attend, we extended an invitation to another candidate with similar
skills/experience identified through online research and word of mouth. The final workshop
included 40 participants.
Establishing open dialogue
Participant privacy was an important consideration throughout workshop planning, and
Institutional Review Board (IRB) approval governed data collection and security for the
workshop. The organizing committee further wished to establish a neutral environment in which
participants holding diverse beliefs and backgrounds would feel comfortable engaging. It was
very important that those attending the workshop felt comfortable sharing their thoughts and
ideas without being concerned about what others might say or do. The planning committee also
decided not to publicize the workshop online beforehand to limit outside attention and encourage
comfort and open discourse among an intimate group of participants. Participants were urged to
avoid taking photos or attributing statements to individuals without permission. The organizers
made efforts to accommodate privacy concerns after they identified a final list of attendees. This
included:
● Name tag options: individuals could simply list their first name with no institutional
affiliation;
● Individuals could choose to remove themselves from some sessions or conversations if
they felt uncomfortable engaging in various topics;
● Photographing other attendees was not permitted unless an attendee received consent
from all individuals who appeared in a photo; and
● Respect for all and approaching conversations with an open mind was a requirement for
participation. If an individual did not feel this was possible, they were asked to not attend.
See email communication sent to all attendees in Appendix B: Guidelines for Conduct.
Workshop Summary
Agenda overview
The event began with a casual, pre-workshop networking social in the evening of August 4,
2025. The organizers provided welcome and opening remarks on the morning of August 5,
2025. Brief participant introductions followed these remarks. A keynote address about the
importance of good UAP data primed participants for the first breakout session (“Identifying,
accessing, and integrating data sources”), held before breaking for lunch. The afternoon of
— 第 5 页 — 由于预算限制和机构能力所限,外部参与受到限制。潜在参与者是根据其在以下一个或多个领域所展现的专业能力来确定的:AI 与机器学习;UAP 研究与数据;物理与自然科学;信息与数据科学;档案与记录;分析方法;网络基础设施与计算;以及人文与社会科学。
如果受邀者谢绝出席,我们会向另一位通过在线检索和口碑推荐所确定的、具有相似技能/经验的候选人发出邀请。最终的研讨会包括 40 名参与者。
建立开放对话
参与者隐私是研讨会筹划全程的一项重要考量,机构审查委员会(Institutional Review Board, IRB)的批准规范了研讨会的数据收集与安全。组织委员会还希望营造一个中立的环境,使持有不同信念和背景的参与者能够自在地参与。让与会者能够在不担心他人言行的情况下自在地分享其想法和观点,这一点非常重要。筹划委员会还决定不事先在网上公开宣传研讨会,以限制外部关注,并在一个亲密的参与者群体中促进舒适与开放的交流。参与者被敦促避免在未经许可的情况下拍照或将言论归于特定个人。组织者在确定最终与会名单后,努力照顾隐私顾虑。这包括:
● 姓名牌选项:个人可以只列出名字,不注明所属机构;
● 如果个人对参与各种话题感到不适,可以选择退出某些环节或对话;
● 未经照片中所有出现者的同意,不允许拍摄其他与会者;以及
● 尊重所有人并以开放的心态对待对话,是参与的一项要求。如果个人认为无法做到这一点,则被要求不要出席。
请参见附录 B:行为准则中发送给所有与会者的电子邮件通信。
研讨会综述
议程概览
活动始于 2025 年 8 月 4 日晚间一场随意的、研讨会前的社交联谊。组织者于 2025 年 8 月 5 日上午致欢迎辞和开幕辞。随后是简短的参与者自我介绍。一场关于优质 UAP 数据重要性的主旨演讲,为参与者的第一场分组讨论(“识别、访问和整合数据来源”)作了铺垫,该讨论在午餐休息前举行。当日下午
— p. 6 — 5
August 5, 2025 began with a plenary talk, followed by the first panel discussion, “Opportunities
and challenges with AI”, and a second breakout session (“Pathways for data analysis and
interpretation at scale”). Day 1 concluded with a brief whole group discussion. A workshop
dinner was held at a restaurant near the workshop venue. Day 2 began with a second plenary talk
and second panel discussion, “Harmonizing qualitative and quantitative perspectives on narrative
data.” After lunch, a series of lightning talks were delivered by participants ahead of the final
breakout session (“Cleaning, organizing and linking data: What can and should be done?”).
Throughout the event, the organizing team collected notes that were later transcribed and
anonymized. For each breakout session, moderators collected records, and notetakers were
assigned to further ensure a robust record of the workshop proceedings.
Breakout Discussion Summaries
Prompts for each breakout session are included in Appendix D: Breakout Session Prompts.
Breakout Session #1: Identifying, accessing, and integrating data sources [DAY 1]
The first breakout session addressed central challenges of UAP research. Discussions revealed
the scope of the UAP data landscape as a patchwork of historical case files, contemporary
narrative reports, sensor-based data (radar, imagery, flight data), and environmental or contextual
datasets (weather, astronomical, seismological). Participants expressed enthusiasm for the
potential to link these disparate sources, but they also acknowledged the barriers posed by
inconsistency in metadata, classification restrictions, missing or inaccessible records, and stigma
around UAP reporting. Despite these challenges, groups converged on the outlook that with clear
standards, prototype integration projects, and intentional collaboration across organizations, it is
possible to create interoperable and sharable datasets that would enable more rigorous and
scalable analysis of UAP reports.
Breakout Session #2: Pathways for data analysis and interpretation at scale [DAY 1]
The second breakout session explored methods and limitations for analyzing UAP narrative data.
Across groups, participants grappled with the tension between extracting operationally useful
signals and respecting the experiential, cultural, and historical richness embedded in reports.
Overall, groups agreed that UAP narratives cannot be reduced to a single analytic approach.
Corpus-level methods (time/space clustering, keyword trends, statistical correlation, graph
analysis) are useful for pattern detection and hypothesis generation, while narrative/experiential
methods (phenomenology, discourse analysis) are useful for preserving meaning, cultural
context, and witness voices. Infrastructures should allow these modes to coexist.
Breakout Session #3: Cleaning, organizing, and linking data: What can and should be
done? [DAY 2]
The third and final breakout activity analyzed the structure of a hypothetical online reporting
form that has collected 1,000 UAP reports stored as PDF files to identify possibilities for data
— 第 6 页 — 即 2025 年 8 月 5 日下午,以一场全体大会演讲开始,随后是第一场专题小组讨论“AI 的机遇与挑战”,以及第二场分组讨论(“大规模数据分析与解读的路径”)。第 1 天以一场简短的全体讨论结束。研讨会晚宴在会场附近的一家餐厅举行。第 2 天以第二场全体大会演讲和第二场专题小组讨论“协调叙述性数据的定性与定量视角”开始。午餐后,参与者作了一系列闪电演讲,之后进入最后一场分组讨论(“清理、组织与关联数据:可以做什么、应该做什么?”)。在整个活动中,组织团队收集了笔记,这些笔记后来被转录并匿名化。对于每场分组讨论,主持人收集记录,并指派记录员以进一步确保对研讨会进程的稳健记录。
分组讨论综述
每场分组讨论的提示载于附录 D:分组讨论提示。
分组讨论 #1:识别、访问和整合数据来源【第 1 天】
第一场分组讨论探讨了 UAP 研究的核心挑战。讨论揭示了 UAP 数据格局的范围,它是历史案卷、当代叙述性报告、基于传感器的数据(雷达、影像、飞行数据),以及环境或背景数据集(天气、天文、地震)的拼凑组合。参与者对关联这些不同来源的潜力表示热情,但也承认了元数据不一致、保密限制、记录缺失或无法访问,以及 UAP 报告污名等所带来的障碍。尽管存在这些挑战,各组一致认为:凭借明确的标准、原型整合项目,以及跨组织的有意协作,有可能创建可互操作、可共享的数据集,从而对 UAP 报告进行更严格、更可扩展的分析。
分组讨论 #2:大规模数据分析与解读的路径【第 1 天】
第二场分组讨论探讨了分析 UAP 叙述性数据的方法与局限。各组参与者努力应对以下张力:既要提取具有作战价值的信号,又要尊重报告中所蕴含的体验性、文化性和历史性的丰富内涵。总体而言,各组一致认为 UAP 叙述不能被简化为单一的分析方法。语料库层面的方法(时间/空间聚类、关键词趋势、统计相关、图分析)有助于模式检测和假设生成,而叙述/体验性方法(现象学、话语分析)则有助于保留意义、文化背景和目击者的声音。基础设施应允许这些模式共存。
分组讨论 #3:清理、组织与关联数据:可以做什么、应该做什么?【第 2 天】
第三场也是最后一场分组活动,分析了一份假想的在线报告表单的结构——该表单已收集了 1,000 份以 PDF 文件形式存储的 UAP 报告——以识别利用所收集数据进行数据
— p. 7 — 6
analysis with the data collected, as well as potential improvement of the form. The discussion led
to the following overarching suggestions that are broadly informative for online UAP reporting
tools.
1. Intake flow and structure:
• Begin with a free-text box (and optional audio upload) where the witness provides their
account in their own words. Use AI-assisted extraction to propose structured fields,
which the witness can then confirm or correct.
• Frame questions around what was perceived (angular size, shape, movement, sound,
effects) rather than presumed properties (exact distance, solid object dimensions).
2. Additions to the form:
• Ask witnesses to explain how they estimated size, distance, or speed (i.e. context
prompts).
• Capture whether this has happened before and, if so, how often.
• Instead of “mass sighting: yes/no”, include approximate numbers of witnesses.
• Include a field for whether the object seemed to react to observer presence.
• Add examples of technological effects (e.g., radio static, car failure) and basic prompts
about feelings or aftereffects that could be informative (e.g., “Did you discuss this with
others? Would you want professional/peer support?”).
• Automatically ingest and display photo metadata (camera model, timestamp, location),
giving users the option to redact sensitive fields.
3. Standardization and cleaning:
• Accept location information including city/address/zip/lat–long, with simple guidance
and drop-downs, and normalize on the back end.
• Enforce a single-entry format for dates and times (calendar widget or drop-downs).
• Allow multiple inputs for units (imperial/metric) but convert and store consistently.
• Include structured numeric fields for object count and multiple objects, with adaptive
follow-up to describe each object separately.
4. Taxonomical considerations:
• Provide a concise taxonomy of common shapes (disk, sphere, triangle, cigar, “other”) but
allow free-text for unusual forms.
• Update descriptive references for cultural familiarity (using objects such as coins or debit
card to estimate size) and internationalize/translate forms for broader accessibility.
5. Integration and linkage:
• Include a field to indicate whether the event was reported elsewhere (NUFORC,
MUFON, FAA, etc.).
— 第 7 页 — 分析的可能性,以及对该表单进行改进的潜力。讨论得出了以下总体建议,这些建议对在线 UAP 报告工具具有广泛的参考价值。
1. 录入流程与结构:
• 以一个自由文本框(以及可选的音频上传)开始,让目击者用自己的话陈述其经历。使用 AI 辅助提取来提议结构化字段,目击者随后可确认或更正这些字段。
• 围绕所感知到的内容(角尺寸、形状、运动、声音、影响)而非假定的属性(确切距离、实体物体尺寸)来设计问题。
2. 表单的补充:
• 请目击者解释他们如何估计尺寸、距离或速度(即背景提示)。
• 记录此前是否发生过此类情况,如有,频率如何。
• 用大致的目击者人数来代替“群体目击:是/否”。
• 包含一个字段,用于记录该物体是否似乎对观察者的存在作出反应。
• 添加技术性影响的示例(例如无线电静电、汽车故障),以及关于可能具有参考价值的感受或后续影响的基本提示(例如“你是否与他人讨论过此事?你是否希望获得专业/同伴支持?”)。
• 自动摄取并显示照片元数据(相机型号、时间戳、位置),并给予用户涂黑敏感字段的选项。
3. 标准化与清理:
• 接受包括城市/地址/邮编/经纬度在内的位置信息,配以简单指引和下拉菜单,并在后端进行规范化。
• 对日期和时间强制采用单一录入格式(日历控件或下拉菜单)。
• 允许多种单位输入(英制/公制),但统一转换和存储。
• 为物体数量和多个物体设置结构化的数字字段,并配以自适应的后续追问以分别描述每个物体。
4. 分类学考量:
• 提供一份简洁的常见形状分类(碟形、球形、三角形、雪茄形、“其他”),但允许对不寻常的形态使用自由文本。
• 更新描述性参照物以贴合文化熟悉度(例如用硬币或借记卡等物体来估计尺寸),并对表单进行国际化/翻译,以扩大可及性。
5. 整合与关联:
• 包含一个字段,用于指明该事件是否曾在别处报告过(NUFORC、MUFON、FAA 等)。
— p. 8 — 7
• Design the schema so reports can be linked to FAA/NASA Aviation Safety Reporting
System (ASRS) data, Automatic Dependent Surveillance-Broadcast (ADS-B) flight
tracks, weather radar, astronomical databases, fireball networks, etc.
• Enable dynamic follow-ups for multiple objects, multiple witnesses, or sequential events.
6. Governance and trust:
• Give reporters clear control over what information (such as geolocation, photo metadata)
is shared publicly.
• Commit to aggregated, de-identified data releases (maps, trend summaries) to build trust
without encouraging hoaxes.
• Light-touch well-being questions were suggested, to help identify if respondents would
like professional or peer follow-up without stepping into clinical assessment.
Outcomes and Recommendations
Synthesis of Findings
Relevant data types and sources of UAP narrative reports
Participants emphasized that UAP research requires drawing on a diverse ecosystem of data,
extending beyond witness testimony. Primary narrative reports in formats ranging from PDFs
and CSVs to emails and oral histories remain central, offering firsthand accounts that, when
digitized and transcribed as needed, can be structured for analysis. These reports are
complemented by smartphone photos and videos, which are widely available but often of poor
quality, though improving over time.
Government sources are handling both classified and unclassified records, including finished
intelligence and historic documents. Military reports and ship logs are particularly robust,
providing structured information on platforms, flight plans, and pilots, while the FAA continues
to collect pilot reports. Other data streams include social media posts, which are often
multimodal (such as online and social media videos); international partner databases; and
structured technical or scientific sensor data, such as radar or spectrum analyses. Supplementary
contextual data is also critical, including flight and weather records, seismological data, satellite
imagery, and even doorbell videos or CCTV systems can corroborate sightings.
Barriers and challenges in data collection and use
Despite many potential sources of data, significant obstacles remain. Access to social media data
has become more restricted due to corporate licensing policies, while ethical and jurisdictional
considerations complicate usage. Classification remains a dominant barrier, as substantial UAP
data may be captured on classified sensors, automatically rendering it inaccessible until
declassified. Other challenges include language and translation barriers, with both human and
automated systems prone to errors, especially in low-resource languages. Stigma in reporting,
— 第 8 页 — • 设计数据结构,使报告能够与 FAA/NASA 航空安全报告系统(Aviation Safety Reporting System, ASRS)数据、广播式自动相关监视(Automatic Dependent Surveillance-Broadcast, ADS-B)飞行航迹、气象雷达、天文数据库、火流星网络等相关联。
• 对多个物体、多名目击者或连续事件启用动态追问。
6. 治理与信任:
• 让报告者对哪些信息(如地理位置、照片元数据)被公开共享拥有明确的控制权。
• 承诺发布经聚合、去标识化的数据(地图、趋势摘要),以在不助长骗局的情况下建立信任。
• 建议设置轻触式的健康福祉问题,以帮助识别受访者是否希望获得专业或同伴的后续跟进,同时不涉入临床评估。
成果与建议
发现的综合
UAP 叙述性报告的相关数据类型与来源
参与者强调,UAP 研究需要借助一个多样化的数据生态系统,其范围超越目击者证词。格式从 PDF、CSV 到电子邮件和口述史的第一手叙述性报告仍是核心,它们提供第一手陈述,在按需数字化和转录后,可被结构化以供分析。这些报告还辅以智能手机拍摄的照片和视频,此类内容虽广泛可得,但往往质量较差,不过正随时间改善。
政府来源正在处理机密和非机密记录,包括成品情报和历史文件。军事报告和舰船日志尤为稳健,提供关于平台、飞行计划和飞行员的结构化信息,而 FAA 则持续收集飞行员报告。其他数据流包括往往是多模态的社交媒体帖子(如在线和社交媒体视频);国际伙伴数据库;以及结构化的技术或科学传感器数据,例如雷达或频谱分析。补充性的背景数据同样至关重要,包括飞行和气象记录、地震数据、卫星影像,甚至门铃视频或闭路电视(CCTV)系统都能佐证目击。
数据收集与使用中的障碍与挑战
尽管有许多潜在的数据来源,重大障碍依然存在。由于企业的许可政策,获取社交媒体数据已变得更加受限,而伦理和管辖权方面的考量则使使用变得复杂。保密仍是一个主要障碍,因为大量 UAP 数据可能是由机密传感器捕获的,在解密之前会自动变得无法访问。其他挑战包括语言和翻译障碍,人工和自动系统都容易出错,在资源匮乏的语言中尤为如此。报告中的污名,
— p. 9 — 8
particularly among pilots, undermines data timeliness and completeness, while the lack of
standardized reporting formats across agencies and organizations further fragments the
landscape. Time sensitivity and weak retention policies have led to the loss of critical records, as
in the well-known Nimitz case. Technical issues are also substantial. Older data can be difficult
to digitize, cursive writing resists Optical Character Recognition (OCR) systems, and
crowdsourced transcription projects suffer from low-quality outputs, recently worsened by
misuse of generative AI. Finally, the field must grapple with fake data and disinformation,
including AI-generated photos or videos, which pose risks for both public trust and analytic
integrity.
Metadata and context for usability and analysis
Effective use of UAP data requires rich contextual metadata. Every report should ideally contain
time, date, and location, preferably with geospatial precision. Distinguishing between descriptive
metadata (objective characteristics like morphology or frequency band) and interpretive metadata
(subjective effects or experiential meaning) is key. Metadata should also capture event-specific
details, such as behaviors, sensor positions, and witness background, and must extend to
technical parameters for structured data. Provenance (the chain of custody and source of the
data) is essential for ensuring interpretability and trust. For visual evidence, metadata such as
device type and embedded geotags allow validation against reported facts. Participants also
emphasized flexible and well-designed reporting forms, for example including “refuse to
answer” options to prevent fabricated entries when respondents lack knowledge.
Linking data sources and developing a unified approach
Given the fragmented nature of UAP data, participants argued for modest, pilot-scale integration
projects as a starting point. Establishing common terminology and data dictionaries is important
to harmonize datasets across agencies and disciplines. Modular and extensible metadata
standards could lead toward a composable ecosystem, potentially implemented through
standardized templates, Interface Control Documents (ICDs), or APIs. Some form of established
governance is needed to facilitate data management and access and engagement for researchers
while alleviating inter-agency silos. Transparency was highlighted as both a goal and a
challenge, as unclassified data should be made available to academia, while sensitive material
must remain protected. Lessons from other fields, such as genetics and astronomy, were cited as
models for developing interoperable metadata standards and ontology-driven approaches.
Assessing credibility and quality of reports
Participants highlighted the importance of sensor reliability, noting that human perception is
fallible. Establishing gold standard exemplars of high-quality reports could help guide future
collection and analysis. Semi-automated triage, assisted by AI, offers promise for sifting through
massive datasets to identify cases with likely conventional explanations as well as cases of
potential interest, though human oversight remains indispensable. Furthermore, credibility is
enhanced when reports are corroborated by multiple witnesses or independent data streams, such
— 第 9 页 — 尤其是在飞行员中,损害了数据的时效性和完整性,而各机构和组织之间缺乏标准化的报告格式,进一步使这一格局碎片化。时间敏感性和薄弱的留存政策已导致关键记录的丢失,正如众所周知的尼米兹(Nimitz)案例。技术问题也相当严重。较旧的数据可能难以数字化,草书难以被光学字符识别(OCR)系统识别,而众包转录项目的产出质量低下,近来因生成式 AI 的滥用而进一步恶化。最后,该领域还必须应对虚假数据和虚假信息,包括 AI 生成的照片或视频,这些对公众信任和分析完整性都构成风险。
用于可用性与分析的元数据与背景信息
有效使用 UAP 数据需要丰富的背景元数据。每份报告理想情况下都应包含时间、日期和地点,最好具有地理空间精度。区分描述性元数据(如形态或频段等客观特征)与解释性元数据(主观影响或体验意义)是关键。元数据还应捕捉事件特定的细节,如行为、传感器位置和目击者背景,并且对于结构化数据必须扩展到技术参数。来源信息(数据的保管链和出处)对于确保可解释性和信任至关重要。对于视觉证据,设备类型和嵌入式地理标签等元数据可用于对照所报告的事实进行验证。参与者还强调了灵活且设计良好的报告表单,例如设置“拒绝回答”选项,以防止受访者在不了解情况时填入虚构条目。
关联数据来源并发展统一方法
鉴于 UAP 数据的碎片化特性,参与者主张以适度的、试点规模的整合项目作为起点。建立通用术语和数据字典对于协调跨机构和跨学科的数据集非常重要。模块化、可扩展的元数据标准可以引向一个可组合的生态系统,并有可能通过标准化模板、接口控制文件(Interface Control Documents, ICD)或 API 来实现。需要某种形式的既定治理,以促进数据管理和访问,以及研究人员的参与,同时缓解机构间的孤岛。透明度被强调为既是目标也是挑战,因为非机密数据应向学术界开放,而敏感材料则必须受到保护。会上引用了遗传学和天文学等其他领域的经验,作为发展可互操作元数据标准和本体驱动方法的范例。
评估报告的可信度与质量
参与者强调了传感器可靠性的重要性,并指出人类感知是易犯错的。建立高质量报告的黄金标准范例,有助于指导未来的收集与分析。由 AI 辅助的半自动分流,在筛选海量数据集以识别可能有常规解释的案例以及具有潜在价值的案例方面颇具前景,尽管人工监督仍不可或缺。此外,当报告得到多名目击者或独立数据流(如
— p. 10 — 9
as radar or weather records. Interviews and psychological screening of witnesses was offered as
an example of how to assess motivations and reduce false reports, though it was acknowledged
that this is difficult to implement at scale. At the same time, biases in favor of certain professions
(pilots, police) must be acknowledged due to enhanced observational training and skills. A
phenomenological approach (qualitative analysis of indicators of lived experiences) allowing
patterns to emerge from narrative accounts was recommended as a complement to quantitative
methods, ensuring that unusual but meaningful details are not prematurely excluded.
AI and analytical methods
AI offers opportunities for pattern recognition, hypothesis generation, and efficiency gains in
large-scale text and multimodal data analysis. Techniques such as semantic search, clustering,
and multimodal modeling (for example, combining acoustic and infrared signals) can help
identify anomalies. AI is also valuable for routine tasks, such as extracting dates or locations
from unstructured text, or triaging likely misidentifications. However, there are risks associated
with AI. Hallucination (the generation of convincing but false conclusions) remains a core
concern. AI analysis is only as reliable as the quality of its input, underscoring the “garbage in,
garbage out” principle. Additionally, LLMs are already biased by UFO-related cultural content,
potentially skewing analyses. Small datasets limit the potential for model training, though pre-
trained models may still be repurposed. Best practices involve an iterative human-AI
collaboration, where algorithms provide preliminary analysis that is verified, corrected, and
enriched by human researchers. Ensemble approaches, leveraging multiple models, may reduce
error rates. Overall, tasks must be carefully defined to align AI methods with research goals,
ensuring a balance between qualitative depth and quantitative rigor.
Forward-looking strategy and key considerations
The group emphasized the need for a forward-looking research infrastructure that integrates
proactive data collection, robust metadata standards, and interdisciplinary collaboration. Some
argued for focusing on new, higher-quality data collection while others urged continued
investment in historical data to preserve its potential value. Future infrastructure priorities
include a unified security solution for managing classified and unclassified data, improved
questionnaire design for witness reports, and benchmarking systems to track analytic
performance over time. Importantly, even “low quality” or stigmatized reports should not be
discarded but made available for diverse lines of inquiry and data reuse.
Finally, participants stressed the need for citizen engagement and ethical responsibility. Public
contributors must be incorporated into coherent strategies for data collection and community-
engaged research. At the same time, researchers must remain vigilant about the risks of
disinformation, AI hallucination, and epistemic injustice, ensuring that narratives are respected in
their original form. Balancing transparency with security, and methodological rigor with
openness to the “weird stuff,” will be essential for building a sustainable, credible, and
innovative field of UAP research.
— 第 10 页 — 雷达或气象记录)的印证时,其可信度会增强。对目击者进行访谈和心理筛查,被作为如何评估动机并减少虚假报告的一个示例,尽管人们也承认这在大规模实施上存在困难。同时,由于飞行员、警察等某些职业受过更强的观察训练和技能,必须承认对这些职业存在偏向。会上建议采用现象学方法(对生活体验指标进行定性分析),让模式从叙述性陈述中浮现,作为定量方法的补充,以确保不寻常但有意义的细节不被过早排除。
AI 与分析方法
AI 为大规模文本和多模态数据分析中的模式识别、假设生成和效率提升提供了机会。语义检索、聚类和多模态建模(例如结合声学与红外信号)等技术有助于识别异常。AI 对于常规任务也很有价值,例如从非结构化文本中提取日期或地点,或对可能的误认进行分流。然而,AI 也伴随着风险。幻觉(生成令人信服但虚假的结论)仍是一个核心担忧。AI 分析的可靠程度取决于其输入的质量,这凸显了“垃圾进,垃圾出”的原则。此外,LLM 已经受到与 UFO 相关的文化内容的偏见影响,可能使分析产生偏差。小型数据集限制了模型训练的潜力,尽管预训练模型或许仍可被重新利用。最佳实践涉及一种迭代式的人-AI 协作,即由算法提供初步分析,再由人类研究人员进行验证、更正和丰富。利用多个模型的集成方法,可能降低错误率。总体而言,必须审慎地界定任务,使 AI 方法与研究目标相一致,确保在定性深度与定量严谨之间取得平衡。
前瞻性战略与关键考量
该小组强调需要一个前瞻性的研究基础设施,将主动的数据收集、稳健的元数据标准和跨学科协作整合起来。有人主张聚焦于新的、更高质量的数据收集,而另一些人则敦促持续投入历史数据,以保留其潜在价值。未来基础设施的优先事项包括:一个用于管理机密和非机密数据的统一安全解决方案,改进目击者报告的问卷设计,以及用于长期追踪分析绩效的基准测试系统。重要的是,即使是“低质量”或被污名化的报告也不应被丢弃,而应可供多种探究路线和数据再利用之用。
最后,参与者强调了公民参与和伦理责任的必要性。必须将公众贡献者纳入连贯的数据收集与社群参与式研究策略中。同时,研究人员必须对虚假信息、AI 幻觉和认知不公正的风险保持警惕,确保叙述以其原始形式得到尊重。在透明度与安全之间、在方法论严谨与对“怪异之事”的开放之间取得平衡,对于建立一个可持续、可信且具创新性的 UAP 研究领域至关重要。
— p. 11 — 10
Recommended actionable next steps
To advance the systematic study of UAP reports and maximize the value of narrative data, the
following actions should be prioritized:
• Develop standardized metadata templates. Standardized metadata should capture core
contextual information while enabling interoperability across agencies and research groups.
Crosswalks between existing schemas will help bridge disciplinary and organizational
differences.
• Adopt a hybrid approach where qualitative expertise and human oversight complement
AI. AI methods can be useful for filtering, transcription, and pattern detection, but human-in-
the-loop infrastructures are essential to help to reduce bias, ensure contextual accuracy, and
improve reliability when working with subjective or sparse data.
• Avoid “reinventing the wheel.” Tools and standards can be adapted from other scientific
fields such as genetics, astronomy, and digital archives. This includes modular metadata
standards, APIs, and citizen science models that can scale efficiently.
• Create systems to triage reports. Reports can be categorized based on credibility, richness
of detail, and corroborating data streams for investigative efficiency. Recognize potential
biases in weighting professional/trained sources (pilots, military) while ensuring that diverse
experiences are preserved for future analysis.
• Capture the meaning and context of reports. Social science expertise and methods should
be implemented alongside quantitative approaches to ensure that experiential and narrative
dimensions are not lost to purely quantitative or technical analysis. This may include
structured interviews, oral histories, and phenomenological coding, among others.
• Continue to preserve and digitize historical reports. Collection of high-quality,
new/contemporary data should be prioritized while preserving and drawing insight from
potentially valuable historical records. A dual focus allows for long-term continuity but
avoids paralysis from the complexity of older archives.
• Enhance public reporting portals. Cross-organizational collaboration can help balance
open participation with safeguards against hoaxes and disinformation. Features such as
including both narrative and structured questions, optional metadata, and “refuse to answer”
options can improve data quality while encouraging participation.
• Convene additional workshops and collaborative opportunities. A strong need was
identified to continue shaping governance structures, standardizing practices, and building an
interdisciplinary community of practice.
— 第 11 页 — 建议的可执行后续步骤
为推进对 UAP 报告的系统性研究并最大化叙述性数据的价值,应优先采取以下行动:
• 开发标准化的元数据模板。标准化的元数据应捕捉核心背景信息,同时实现跨机构和研究团队的互操作性。现有数据结构之间的对照映射(crosswalks)将有助于弥合学科和组织间的差异。
• 采用一种混合方法,让定性专业知识和人工监督对 AI 加以补充。AI 方法对于筛选、转录和模式检测很有用,但当处理主观或稀疏数据时,人在回路(human-in-the-loop)的基础设施对于帮助减少偏见、确保背景准确性和提高可靠性是必不可少的。
• 避免“重复造轮子”。可以借鉴遗传学、天文学和数字档案等其他科学领域的工具和标准。这包括可高效扩展的模块化元数据标准、API 和公民科学模式。
• 创建对报告进行分流的系统。可根据可信度、细节丰富程度和相互印证的数据流对报告进行分类,以提高调查效率。要认识到在赋予专业/受训来源(飞行员、军人)更高权重时可能存在的偏见,同时确保多样化的经历得以保存以供未来分析。
• 捕捉报告的意义与背景。社会科学的专业知识和方法应与定量方法一并实施,以确保体验性和叙述性维度不会因纯定量或技术分析而丢失。这可能包括结构化访谈、口述史和现象学编码等。
• 持续保存和数字化历史报告。应优先收集高质量的新/当代数据,同时保存并从潜在有价值的历史记录中汲取洞见。这种双重侧重有助于长期的连续性,但可避免因旧档案的复杂性而陷入瘫痪。
• 增强公众报告门户。跨组织协作有助于在开放参与与防范骗局和虚假信息之间取得平衡。诸如同时包含叙述性和结构化问题、可选元数据以及“拒绝回答”选项等功能,可以在鼓励参与的同时提高数据质量。
• 召开更多的研讨会与协作机会。会上明确指出,强烈需要继续塑造治理结构、标准化实践,并建立一个跨学科的实践社群。
— p. 12 — 11
Appendix A: Invitation Letter
Subject: Invitation to participate in workshop on Unidentified Anomalous Phenomena (UAP)
narrative data integration and analysis
We are writing to invite you to participate in the upcoming “2025 UAP Workshop: Narrative
Data, Infrastructures, and Analysis”, which will be held August 5-6, 2025, in the
Washington, DC area. This in-person, two-day event will bring together experts, researchers,
and stakeholders to discuss best practices for Unidentified Anomalous Phenomena (UAP)
narrative data collection, management, interoperability, and analysis methodologies to enhance
transparency and scientific rigor in UAP research.
Analyzing reports of UAP related sightings and experiences presents unique challenges due to
the large-scale, heterogeneous, and qualitative nature of the reports originating from military and
civilian sources. These reports typically lack standardized metadata, formatting, or nomenclature,
making comparative analysis difficult. Additionally, the integration of UAP reports from
disparate sources—such as military databases, online reporting systems, digital and digitized
archival records, and social media—poses significant challenges for harmonization and
verification of data and construction of evidence. The complexity of these datasets requires
innovative data infrastructure solutions to enhance reliability, accessibility, and interoperability.
The workshop will explore these challenges and seek strategies to improve UAP data
standardization, integration, and analytical approaches.
The workshop builds upon a previous NSF-funded workshop held in 2024, “Unidentified
Anomalous Phenomena (UAP): A Dialogue on Science, Public Engagement and
Communication”. A new collaboration between the All-domain Anomaly Resolution Office
(AARO), Associated Universities, Inc. (AUI), and Florida State University (FSU) has convened
to host the upcoming event on the topic of narrative data.
Given your experience and expertise in areas such as UAP studies, data collection, management,
interoperability, and/or analysis methodologies, we believe that you can make significant
contributions to this gathering and its outcomes. If you choose to participate, you will have an
opportunity to interact and collaborate with a small group of 25-30 experts, including
information and data science researchers and practitioners, and government stakeholders focused
on proposing a roadmap for the design and implementation of tools that will advance science,
improve data management practice, and inform investments for UAP research.
While we are unable to compensate you for travel costs to/from the Washington, DC area, we
will cover lodging expenses and most meals for the duration of the event. We expect this to be a
groundbreaking workshop and meaningful gathering of minds that will help to shape the future
of UAP research.
If you are interested and can attend the event in person, we kindly ask that you fill out this
form by June x, 2025:
— 第 12 页 — 附录 A:邀请函
主题:邀请参加关于不明异常现象(UAP)叙述性数据整合与分析的研讨会
我们谨此邀请您参加即将举行的“2025 UAP 研讨会:叙述性数据、基础设施与分析”,会议将于 2025 年 8 月 5-6 日在华盛顿特区地区举行。这场为期两天的线下活动将汇集专家、研究人员和利益相关方,讨论不明异常现象(UAP)叙述性数据收集、管理、互操作性和分析方法论的最佳实践,以增强 UAP 研究的透明度和科学严谨性。
分析与 UAP 相关的目击和体验报告面临独特的挑战,因为这些源自军方和民间的报告具有大规模、异质和定性的特点。这些报告通常缺乏标准化的元数据、格式或术语,使得比较分析变得困难。此外,整合来自不同来源——如军事数据库、在线报告系统、数字及数字化档案记录、社交媒体——的 UAP 报告,为数据的协调、验证和证据构建带来了重大挑战。这些数据集的复杂性要求创新的数据基础设施解决方案,以增强可靠性、可及性和互操作性。研讨会将探讨这些挑战,并寻求改进 UAP 数据标准化、整合与分析方法的策略。
本次研讨会建立在此前一场于 2024 年举办的、由美国国家科学基金会(NSF)资助的研讨会“不明异常现象(UAP):关于科学、公众参与和传播的对话”之上。全域异常解决办公室(AARO)、联合大学公司(AUI)和佛罗里达州立大学(FSU)之间新建立的一项合作,共同召集主办了这场以叙述性数据为主题的即将举行的活动。
鉴于您在 UAP 研究、数据收集、管理、互操作性和/或分析方法论等领域的经验与专长,我们相信您能为这次聚会及其成果作出重大贡献。如果您选择参与,您将有机会与一个由 25-30 名专家组成的小团体互动与协作,其中包括信息与数据科学的研究人员和实践者,以及专注于为设计和实施推进科学、改善数据管理实践并为 UAP 研究投资提供参考的工具提出路线图的政府利益相关方。
虽然我们无法补偿您往返华盛顿特区地区的差旅费用,但我们将承担活动期间的住宿费用和大部分餐食。我们期待这将是一场开创性的研讨会,以及一次有意义的思想聚会,将有助于塑造 UAP 研究的未来。
如果您有兴趣并能亲自出席活动,我们恳请您在 2025 年 6 月 x 日前填写此表单:
— p. 13 — 12
[Event Participation Form]
If you cannot or do not plan to attend, we’d greatly appreciate it if you could let us know by
responding to this email.
We sincerely hope you accept this invitation and look forward to seeing you at the event. If you
are unable to attend, please feel free to reply and nominate others with similar expertise
who may wish to attend. Note that due to limited budget and capacity, invitations are selective
based on expertise. However, all outcomes will be communicated promptly, publicly, and
transparently.
If you have any questions or require further information, please do not hesitate to contact us.
On behalf of the Organizing Committee,
Gretchen Stahlman, Florida State University School of Information
Tim Spuck, Associated Universities, Inc.
— 第 13 页 — 【活动参与表单】
如果您无法或不打算出席,我们将非常感谢您回复此邮件告知我们。
我们真诚地希望您接受此邀请,并期待在活动上与您见面。如果您无法出席,请随时回复并提名其他可能希望参加的、具有相似专长的人士。请注意,由于预算和容量有限,邀请是基于专长而有选择性的。然而,所有成果都将被及时、公开、透明地传达。
如果您有任何疑问或需要进一步的信息,请随时与我们联系。
谨代表组织委员会,
Gretchen Stahlman,佛罗里达州立大学信息学院
Tim Spuck,联合大学公司
— p. 14 — 13
Appendix B: Guidelines for Conduct
Workshop Focus: Please keep in mind that the workshop focus is UAP data. Specifically,
discussion will focus on:
• Identifying, accessing, and integrating data sources,
• Opportunities and challenges with AI,
• Pathways for data analysis and interpretation at scale,
• Harmonizing qualitative and quantitative perspectives on narrative data,
• Cleaning, organizing, and linking data, and
• What can and should be done to advance UAP research.
It is important to note that we will NOT be focused on what UAP are or their origins, but rather
how we can best acquire, curate, and analyze UAP data (qualitative and quantitative) through the
scientific lens. It is our hope that the work we do together can better normalize the conversation,
data collection, data analysis, etc. about UAP within the research community and the general
public. When we make something “taboo” or become dismissive of others’
observations/experiences, we create an opening for misinformation and misunderstanding to
thrive. That is not good for anyone. Einstein once said, “No amount of experimentation can
prove me right. A single experiment can prove me wrong.” This is why science must remain
exploratory in nature. Just because we believe something to be true, it must not prevent us from
giving consideration to an alternative conclusion supported by evidence.
We fully expect there to be diverse opinions presented throughout the workshop, and we want to
take time to make sure everyone who is attending will be comfortable sharing and will feel their
thoughts and ideas are valued and respected. Our intent is to ensure we are creating a space
where everyone can share what they are comfortable sharing. In an effort to create this space, we
will be asking all participants to abide by these rules:
• During the Workshop, do NOT take pictures or video that include other attendees,
• You may hear someone say something of interest during the workshop. Please secure
their permission before repeating it and attributing it to them to others who are not in
attendance at the workshop,
• While you may disagree with the perspective of others, we ask that everyone work to
maintain a professional and respectful attitude throughout,
• During breakout sessions, we plan to record the discussions and have them transcribed.
Any identifiable data will be removed, and the recordings will be destroyed once
transcription and verification has been completed. In the event that an individual in a
breakout room does not want to be recorded, we will not record the session, but hand-
written notes will be taken throughout the session.
— 第 14 页 — 附录 B:行为准则
研讨会焦点:请记住,研讨会的焦点是 UAP 数据。具体而言,讨论将聚焦于:
• 识别、访问和整合数据来源,
• AI 的机遇与挑战,
• 大规模数据分析与解读的路径,
• 协调叙述性数据的定性与定量视角,
• 清理、组织与关联数据,以及
• 为推进 UAP 研究可以做什么、应该做什么。
需要重点指出的是,我们的关注点并非 UAP 是什么或其来源,而是我们如何能够通过科学的视角,最好地获取、整理和分析 UAP 数据(定性与定量)。我们希望我们共同开展的工作,能够更好地使研究界和公众之间关于 UAP 的对话、数据收集、数据分析等趋于正常化。当我们把某件事视为“禁忌”,或对他人的观察/体验不屑一顾时,我们就为错误信息和误解的滋生创造了可乘之机。这对任何人都没有好处。爱因斯坦曾说:“再多的实验也无法证明我是对的。一个实验就足以证明我是错的。”这正是科学必须保持探索性本质的原因。仅仅因为我们相信某事为真,这不应妨碍我们去考虑一个有证据支持的替代结论。
我们完全预料到整个研讨会中会呈现多样化的意见,我们希望花时间确保每一位与会者都能自在地分享,并感到他们的想法和观点受到重视和尊重。我们的意图是确保我们正在创造一个让每个人都能分享其愿意分享内容的空间。为了创造这一空间,我们将要求所有参与者遵守以下规则:
• 在研讨会期间,请勿拍摄包含其他与会者的照片或视频,
• 您在研讨会期间可能会听到某人说了些有意思的话。在向未出席研讨会的他人复述并将其归于该人之前,请先征得其许可,
• 尽管您可能不同意他人的观点,我们请求每个人都努力在全程保持专业和尊重的态度,
• 在分组讨论期间,我们计划录制讨论并进行转录。任何可识别的数据都将被删除,一旦转录和核验完成,录音将被销毁。如果分组讨论室中的某位个人不希望被录音,我们将不录制该环节,但会在整个环节中做手写笔记。
— p. 15 — 14
Appendix C: Workshop Agenda
August 4th:
Pre-workshop day – 7 pm to 9 pm
• 7:00-9:00 pm: Networking social
o Light dinner near workshop venue
August 5th:
9 am to 5 pm
• 8:00-9:00: Breakfast
• 9:00-9:15: Welcome, announcements, and workshop format overview
• 9:15-9:30: Self-introduction of participants
• 9:30-10:45: Keynote: Setting the stage: Importance of good UAP data
• 10:45-11:00: Coffee break
• 11:00-12:00: Breakout session #1: Identifying, accessing, and integrating data sources
• 12:00-12:15: Report out
• 12:15-1:15: Lunch and networking
• 1:15-2:15: Plenary #1
• 2:15-2:45: Coffee break
• 3:45-4:45: Panel #1
• 4:45-5:45: Breakout session #2: Pathways for data analysis and interpretation at scale
• 5:45-6:00: Report out and wrapping up
• 7:00: Light dinner near workshop venue
August 6th:
9 am to 4 pm
• 8:00-9:00: Breakfast
• 9:00-10:00: Plenary #2
• 10:00-10:30: Participant presentations #1 (lightning talks)
• 10:30-10:45: Coffee break
• 10:45-11:45: Panel #2
• 11:45-1:00: Lunch and networking
• 1:00-2:10: Participant presentations #2 (lightning talks)
• 2:10-2:30: Coffee break
• 2:30-3:30: Breakout session #3: Cleaning, organizing, and linking data: What can and
should be done? (Interactive hackathon)
• 3:30-3:45: Report out
• 3:45-4:00: Whole group discussion and closing remarks
• 4:00: Workshop concludes
— 第 15 页 — 附录 C:研讨会议程
8 月 4 日:
研讨会前一天 — 晚 7 点至 9 点
• 晚 7:00-9:00:社交联谊
o 会场附近的简餐
8 月 5 日:
上午 9 点至下午 5 点
• 8:00-9:00:早餐
• 9:00-9:15:欢迎、通知及研讨会形式概览
• 9:15-9:30:参与者自我介绍
• 9:30-10:45:主旨演讲:铺垫背景——优质 UAP 数据的重要性
• 10:45-11:00:咖啡休息
• 11:00-12:00:分组讨论 #1:识别、访问和整合数据来源
• 12:00-12:15:汇报
• 12:15-1:15:午餐与社交
• 1:15-2:15:全体大会 #1
• 2:15-2:45:咖啡休息
• 3:45-4:45:专题小组 #1
• 4:45-5:45:分组讨论 #2:大规模数据分析与解读的路径
• 5:45-6:00:汇报与收尾
• 7:00:会场附近的简餐
8 月 6 日:
上午 9 点至下午 4 点
• 8:00-9:00:早餐
• 9:00-10:00:全体大会 #2
• 10:00-10:30:参与者展示 #1(闪电演讲)
• 10:30-10:45:咖啡休息
• 10:45-11:45:专题小组 #2
• 11:45-1:00:午餐与社交
• 1:00-2:10:参与者展示 #2(闪电演讲)
• 2:10-2:30:咖啡休息
• 2:30-3:30:分组讨论 #3:清理、组织与关联数据:可以做什么、应该做什么?(互动式黑客松)
• 3:30-3:45:汇报
• 3:45-4:00:全体讨论与闭幕致辞
• 4:00:研讨会结束
— p. 16 — 15
Appendix D: Breakout Session Prompts
Breakout Session #1: Identifying, Accessing, and Integrating Data Sources
Time: Day 1, 11:00–12:00 | Report-out: 12:00–12:15
Purpose: Surface knowledge about existing UAP-related data sources, understand integration
barriers, and identify opportunities for improving data discovery and interoperability.
Group Structure: Form 4 groups (~8–10 people each). Groups will be labeled “1”, “2”, “3”,
and “4” (numbers will be written on the back of name badges). Facilitator will record the session
with audio recorder or cell phone. Facilitator or designated note-taker for each group will take
notes. The group will elect a reporter to report out to the whole workshop following the breakout
session.
Guiding Questions:
• What are the relevant data types and sources of UAP narrative reports?
• What are the key barriers to accessing these diverse datasets?
• What metadata or context is needed to curate and make these sources usable for analysis?
• How can we responsibly link or integrate these diverse data sources?
• What other data sources can be used to corroborate UAP narrative report data?
• How do we assess credibility and quality of the narrative report data?
Breakout Session #2: Pathways for Data Analysis and Interpretation at Scale
Time: Day 1, 4:45–5:45 | Report-out: 5:45–6:00
Purpose: Explore methods for analyzing large volumes of UAP data, and identify pathways for
scalable, reliable, and ethically sound interpretation.
Group Structure: Form 4 groups (~8–10 people each). Groups will be labeled “A”, “B”, “C”,
and “D” (letters will be written on the back of name badges). Facilitator will record the session
with audio recorder or cell phone. Facilitator or designated note-taker for each group will take
notes. The group will elect a reporter to report out to the whole workshop following the breakout
session.
Guiding Questions:
• What analytical methods are best suited to UAP narrative report data?
• How can structured and unstructured data be combined?
• How can we responsibly and ethically identify and interpret patterns in large datasets?
• Where is AI promising, and where is human interpretation essential, and how can the two
work together?
• What should be prioritized for future narrative data analysis infrastructure?
— 第 16 页 — 附录 D:分组讨论提示
分组讨论 #1:识别、访问和整合数据来源
时间:第 1 天,11:00–12:00 | 汇报:12:00–12:15
目的:呈现关于现有 UAP 相关数据来源的知识,理解整合障碍,并识别改进数据发现和互操作性的机会。
分组结构:组成 4 个小组(每组约 8–10 人)。各组将被标记为“1”“2”“3”和“4”(数字将写在姓名牌背面)。主持人将用录音设备或手机录制该环节。每组的主持人或指定记录员将做笔记。小组将推选一名汇报人,在分组讨论后向整个研讨会汇报。
引导性问题:
• UAP 叙述性报告的相关数据类型与来源有哪些?
• 访问这些多样化数据集的关键障碍是什么?
• 需要哪些元数据或背景信息来整理这些来源并使其可用于分析?
• 我们如何负责任地关联或整合这些多样化的数据来源?
• 还有哪些其他数据来源可用于佐证 UAP 叙述性报告数据?
• 我们如何评估叙述性报告数据的可信度与质量?
分组讨论 #2:大规模数据分析与解读的路径
时间:第 1 天,4:45–5:45 | 汇报:5:45–6:00
目的:探索分析大量 UAP 数据的方法,并识别可扩展、可靠且符合伦理的解读路径。
分组结构:组成 4 个小组(每组约 8–10 人)。各组将被标记为“A”“B”“C”和“D”(字母将写在姓名牌背面)。主持人将用录音设备或手机录制该环节。每组的主持人或指定记录员将做笔记。小组将推选一名汇报人,在分组讨论后向整个研讨会汇报。
引导性问题:
• 哪些分析方法最适合 UAP 叙述性报告数据?
• 结构化和非结构化数据如何结合?
• 我们如何负责任且合乎伦理地识别和解读大型数据集中的模式?
• AI 在何处有前景,人类解读在何处必不可少,二者如何协同?
• 未来的叙述性数据分析基础设施应优先考虑什么?
— p. 17 — 16
Breakout Session #3: Pathways for Data Analysis and Interpretation at Scale
Time: Day 2, 2:30-3:30 | Report-out: 3:30-3:45
Purpose: Surface knowledge about existing UAP-related data sources, understand integration
barriers, and identify opportunities for improving data discovery and interoperability.
Group Structure: Form 4 groups (~8–10 people each). Facilitator will record the session with
audio recorder or cell phone. Facilitator or designated note-taker for each group will take notes.
The group will elect a reporter to report out to the whole workshop following the breakout
session. A supplementary printed document will be provided to all participants for review
showing data fields collected by the reporting tool described in the scenario below.
Scenario: An agency has launched a new online tool for the public to report UAP. The tool has
collected over 1,000 reports, each saved as a PDF (see fields in printed document).
Guiding Questions:
1. What would you change about the data collection form?
2. What could you do with this dataset?
3. What cleaning techniques can be applied?
How could this dataset be linked to others, such as data from other non-government
organizations?
— 第 17 页 — 分组讨论 #3:大规模数据分析与解读的路径
时间:第 2 天,2:30-3:30 | 汇报:3:30-3:45
目的:呈现关于现有 UAP 相关数据来源的知识,理解整合障碍,并识别改进数据发现和互操作性的机会。
分组结构:组成 4 个小组(每组约 8–10 人)。主持人将用录音设备或手机录制该环节。每组的主持人或指定记录员将做笔记。小组将推选一名汇报人,在分组讨论后向整个研讨会汇报。将向所有参与者提供一份补充打印文件供审阅,其中展示下述情景中所述报告工具所收集的数据字段。
情景:某机构推出了一个供公众报告 UAP 的新在线工具。该工具已收集了超过 1,000 份报告,每份都保存为一个 PDF(见打印文件中的字段)。
引导性问题:
1. 关于该数据收集表单,您会作出哪些更改?
2. 您能用这个数据集做什么?
3. 可以应用哪些清理技术?
这个数据集如何能与其他数据集关联,例如来自其他非政府组织的数据?