AI Research for National Security
With AI rapidly reshaping the geopolitical landscape, research scholars, intelligence analysts, and software developers gathered at NC State for LAS’s fifth annual Summer Conference on Applied Data Science to safely explore the application of new tech for national security.
At the front of a classroom in Fitts-Woolard Hall at NC State University, researcher Stephen S. slowly pulls a shiny, gold-chrome bust from a canvas bag and places it on the table beside him. The life-sized head-and-shoulders toy sculpture is of the famous humanoid robot from “Star Wars.” It’s the first day of the Summer Conference on Applied Data Science, and Stephen is explaining the very recent evolution of AI to attendees.
“Unlike C-3PO, who just takes direction and does what he’s told, an agentic AI makes decisions, like R2-D2,” Stephen tells the group, unveiling a much smaller figurine of the film’s other beloved droid.
Turning the vision of autonomous technology into practical tools for national security is a core focus of the summer-long research program known as SCADS. The intensive eight-week workshop organized by the Laboratory for Analytic Sciences (LAS) at NC State aims to address data challenges faced by the U.S. Intelligence Community. A group of 27 government, industry, and academic collaborators with unique perspectives on a select set of problems gathered daily on Centennial Campus to build AI tools for decision-makers working in defense and national security.
The 2026 SCADS and its associated problem book tackled how agentic AI can help intelligence analysts make sense of massive datasets. To ensure a successful technology transfer, researchers must carefully evaluate the security and governance risks of these autonomous tools. Ultimately, these systems are only useful for intelligence analysts if they work on the “high side” – the government’s classified environment where confidential and top-secret information is shared – making performance evaluation and implementation on government resources the conference’s foundational goals.
“We have to build a bridge that allows us to bring these AI tools to resource-constrained environments, like the secure, classified networks used by government staff,” Stephen says.
Most collaborators from academia and industry don’t hold security clearances, so the conference also has a secondary goal of creating or augmenting synthetic datasets to simulate real-world data when the actual data is classified and cannot be shared.
The Rise of the AI Agent
Now in its fifth year, SCADS has grown up with AI. During the summer of 2022, most inaugural SCADS participants were familiar with DALL-E 2, an AI system that could generate images from text. Just a few months later, ChatGPT’s game-changing, user-friendly interface launched, bringing generative AI to millions of people around the world in 2023. In 2024, AI became more multimodal, processing not just text but also audio, video, and images. By 2025, users could tell AI to reason and think before answering.
But software development has changed more in the past six months than it has in the past 25 years, say staff at the lab. The focus this year was on agentic AI – technology that can take actions autonomously on behalf of a person or institution to carry out multistep processes without explicit, step-by-step instructions, and can efficiently adapt as it goes.
“This year, we’ve seen [AI] write entire software packages if you give it the right instruction,” says Stephen S. “This year has been revolutionary.”

The intelligence community believes it can make more informed decisions when armed with the insight of those in academia and industry working on similar challenges. This cross-sector collaboration is what makes LAS unique. Beyond local experts from NC State, the University of North Carolina at Charlotte, Winston-Salem State University, IBM, and SAS, this year’s SCADS participants and visitors hailed from AWS, Rochester Institute of Technology, Tufts University, Northeastern University, University of Georgia, Georgia Tech, University of Texas – San Antonio, Virginia Tech, FBI, U.S. Air Force, Pacific Northwest National Laboratory, Institute for Defense Analysis, West Point Military Academy, and the NSA.
Research conducted at SCADS aligns with LAS’s ongoing work in human-centered AI, audio and video sensemaking, and operationalizing AI and ML. To guide participant efforts, Stephen S. and fellow technical staff at the lab collaborated with stakeholders across the U.S. Intelligence Community to define the SCADS problem book. Initially detailing nine core projects for cross-sector collaborative teams, the initiative expanded to 11 distinct projects presented by attendees at the conclusion of the conference.
Research on Unclassified Datasets
Working with publicly available data that shares characteristics with classified data helps academic researchers and software developers without an active security clearance build better tools for government analysts. SCADS participants used datasets such as a collection of millions of complaint calls from New York City neighborhoods, an enormous database of user manuals for antique vehicles and power tools, and a collection of emails released after the Enron scandal to test their ideas. Some teams also created their own datasets.
Playing Telephone in Reverse: Fixing Transcription Errors
Automatic speech recognition systems convert spoken words into written text – for example, dictating voice-to-text messages on our phones. Still, they struggle with audio that contains background noise or overlapping speech, often garbling or missing names, numbers, and other terms that make a transcript searchable and accurate. One project at SCADS involved improving transcription accuracy by using additional context. Researchers improved the performance of a speech-to-text AI model by accessing additional data related to the topics discussed.
To test their approaches, the team working on the Playing Telephone in Reverse project built their own dataset. They wrote conversation scripts in different languages, then recorded themselves reading them under challenging acoustic conditions, such as in a restaurant, in a moving car with music playing, on a windy day, and from a distance from the microphone.
The group was able to determine optimal ways to combine multiple language models and allow an AI model with context to provide reconciliation of the differences in the outputs. This boosted the success rate from 68% to 80%.

Missing Accomplice: Creating an Org Chart from the Enron Emails
In a large communication archive, the organization’s structure is buried in the traffic rather than stated alongside it. Analysts need to know who reports to whom and how formal authority maps onto informal influence, but that has to be inferred from the data, then navigated at scales where existing graph tools break down.
A SCADS team built Missing Accomplice, a pattern-discovery dashboard with a tool-first, LLM-second architecture: the language model orchestrates deterministic graph queries and returns every finding with the underlying records attached, rather than asserting conclusions as fact. Analyst corrections and working hypotheses are captured as a reviewable overlay rather than as silent edits to the data. The team demonstrated the system end-to-end on the Enron Corporation email dataset, which contains about half a million internal emails from roughly 150 people. It was made public after the company’s 2001 financial collapse from accounting fraud.
Ashley Anderson, a professor at Virginia Tech specializing in human-centered design, worked with other SCADS participants from the intelligence community on the interface. In one test run, she used the tool to trace a specific thread through the Enron data to determine whether the company’s fraud cover-up was being done by an organized group of employees.
“Our premise is that fraud at this scale requires coordination,” Anderson says. “We’re showing how this system finds a tight, highly active inner circle that was managing Enron’s narrative. The coordination is hidden in the structure of communications, not just their content.”
The team also tested the scalability of their dashboard on an even larger dataset containing 2.3 million emails from 82,000 email addresses, yielding positive results.
Used against classified data, the same approach could help intelligence analysts map roles, reporting lines, and informal influence across a network of bad actors.

A Collaborative Mindset and Shared Mission
NC State is one of the nation’s leading universities for military research, with Department of Defense-funded projects totaling $49 million in annual research expenditures in fiscal year 2024, according to the National Science Foundation. The hands-on problem-solving at SCADS embodies the core of the university’s Think and Do attitude. As the summer participants return to their schools and jobs, staff at the Laboratory for Analytic Sciences are working with government stakeholders to deploy these new AI tools for national security decision-makers.
“They’ve taken steps toward designing the intelligence analysis system of the future,” says Stephen S.
These agentic systems will not replace human judgment, but will serve as reliable, decision-making “R2-D2” partners to the nation’s intelligence analysts.
- Categories: