A research team studied how humans and AI agents collaborated to build a new AI model. The project examined expectations about how independently agents can function. Researchers from China's Fudan University analyzed their own project to understand the division of labor. They reviewed over 700 task logs from 56 participants alongside logs from the AI agents.
The project centered on developing an agentic language model named Atria Dawn Preview. The system uses a mixture-of-experts architecture with 744 billion parameters. It is designed specifically for research and engineering tasks. Training relied on a pipeline tying each task to a real execution environment. The model calls tools, generates intermediate results, and gets checked against external signals.
The team reports that the model leads on five of 16 benchmarks. These include web search and cybersecurity tests. However, the model does not hold an overall edge over its competitors. AI tools were utilized in 96.5 percent of the reviewed tasks during the project.
Read nextAI Agents Grow Autonomous While Enterprises Struggle with Oversight and ControlParticipants handed off more work over time
Participants handed off more work to agents over the course of four weeks. The median ratio of agent actions to human inputs climbed from 11 to 28.5. The study team cautions against reading this trend as growing autonomy. Each human decision led to more agent steps rather than independent agent choices.
Participants reported that roughly a third of completed AI-assisted tasks were infeasible without AI. Out of 455 completed tasks, 151 required AI assistance to reach the same scope and quality. These tasks spread across 27 of the 56 participants. AI made work possible that would never have been started otherwise.
Humans consistently made the final choices across all tested scenarios. For methods and parameters, the pattern of AI proposing and humans selecting occurred 55.4 percent of the time. Humans made 85.5 percent of decisions about methods and parameters. AI made just 9.2 percent of those choices.

Humans made almost all final decisions
Humans finalized goals and scope in 93.4 percent of cases. AI proposal shares ranged from 17 to 55 percent depending on decision types. Final decision shares for AI remained in the single digits. Humans chose the goal 95.4 percent of the time even in tasks rated infeasible without AI.
Human intervention moved work forward in 76 percent of problem cases. Agents solved problems independently in only 23 percent of those instances. Human help typically involved supplying context or diagnosing issues. When AI outputs needed revision, the AI handled changes itself 75.4 percent of the time after receiving feedback.
The authors warn of a rubber-stamp risk when oversight becomes difficult over long chains of agent work. Many participants ran agents autonomously to avoid constant approvals. Anthropic considers an AI that develops its successor possible sooner than expected. OpenAI uses GPT-5.6 Sol across its development cycle, while Google and DeepMind use Dream-RSI.



