Designing Evidence-Based AI Pilots for Higher-Education Learning

This episode examines how universities are moving from broad AI enthusiasm toward bounded pilots, instructor oversight, and evidence-based implementation. It considers AI’s potential to support practice and feedback while protecting independent thinking, productive struggle, and knowledge development.

[Jordan]: Welcome to AI in Higher Education. I’m Jordan. Today, we’re looking at a shift from broad AI enthusiasm toward more deliberate experimentation: where these tools genuinely support learning, where they can get in the way, and how institutions are building evidence before scaling up.

[Jordan]: Let’s begin with a useful question from a new University of Minnesota report. Not simply, “Is AI good or bad for education?” but, “Which kind of learning is AI helping—or replacing?”

[Jordan]: The project brought together 30 learning and education specialists to examine 65 different learning processes. Their strongest findings focused on practice, feedback, scaffolding, self-testing, and helping students consolidate what they understand.

[Jordan]: But the risks were just as important. AI can undermine initial knowledge acquisition, higher-order thinking, metacognition, self-regulated learning, transfer, and social learning when it performs the cognitive work students need to do themselves.

[Jordan]: The experts recommend a sequence that will sound familiar to anyone involved in good instructional design. Students first learn without AI. Then, they use AI to strengthen their understanding. Finally, they demonstrate independent thinking in new contexts.

[Jordan]: That matters because it gives faculty something more practical than a blanket policy. The question becomes: what learning process is this activity supposed to strengthen? And which parts of the work must remain deliberately human and effortful?

[Jordan]: The report also cautions that much of the existing evidence is still methodologically weak. So even as institutions act, they need to stay modest about what they think they know.

[Jordan]: That idea of bounded experimentation appears clearly in the University of Iowa’s current pilots.

[Jordan]: Iowa is recruiting instructors to test course-integrated AI tools, including a general Class Tutor and customizable, scenario-based assistants. Faculty can review chat logs, which creates visibility into how students are using the tools.

[Jordan]: A second pilot is looking at AI-assisted grading and feedback based on instructor-created rubrics. The use cases include essays, audio and video presentations, and spreadsheets—but with human oversight.

[Jordan]: What’s notable here is that AI is being treated as an instructional intervention, not just as a general-purpose chatbot dropped into a course.

[Jordan]: The pilots ask more specific questions. Can scenario-based practice help students learn? Can rubric-grounded feedback be consistent and useful? Can instructors remain accountable for academic judgment?

[Jordan]: Those are much better questions than simply asking whether students like the tool, or how often they use it.

[Jordan]: And Iowa says it will support setup, course integration, and evaluation. That operational support matters. A promising tool can still produce weak results if faculty are left to figure out the pedagogy, configuration, and evidence-gathering on their own.

[Jordan]: A different development is coming from Marshall University. Marshall Online has received a $200,000 Education to Workforce Impact Fund grant for a project called “Teaching AI How We Teach.”

[Jordan]: The project will create an AI-enabled skills-to-knowledge graph and an implementation-intelligence agent. The goal is to examine how skills developed in courses correspond to workforce needs.

[Jordan]: The initial pilot includes public administration, communication disorders, cybersecurity, and pharmaceutical sciences, with possible expansion to other programs and institutions.

[Jordan]: This represents a shift in how colleges are thinking about AI. Instead of using it mainly to generate course content, the project is using AI to make curricular knowledge and student competencies more visible across academic and employment settings.

[Jordan]: If it works well, that could help programs identify skill gaps, improve course-to-career mapping, and offer clearer evidence of what students can actually do.

[Jordan]: But there’s an important condition. Inferred skills and labor-market alignment still have to be validated. An automated system shouldn’t turn a complicated educational judgment into an unquestioned fact.

[Jordan]: That makes this especially relevant to academic program leaders, career services, workforce partnerships, and accreditation teams. The promise is real, but so is the need for careful review.

[Jordan]: Stanford offers another model—one focused on supporting many small experiments, rather than waiting for a single, institution-wide AI strategy.

[Jordan]: Stanford’s AI Meets Education initiative received more than 60 proposals and funded 12 course and curriculum projects across four schools and units. Together, those projects affect 25 course offerings and an estimated 3,700 students.

[Jordan]: The initiative’s early reflections highlight a concern many instructors share: students may use AI to avoid productive struggle.

[Jordan]: The funded projects are exploring new approaches to course design, assessment, and learning in an AI-rich environment. The broader lesson is organizational. Universities may learn more by funding multiple bounded projects, providing educational-design support, and looking for patterns across them.

[Jordan]: In other words, the goal isn’t to find one perfect AI policy for every discipline. It’s to create a structure in which faculty can test meaningful changes and share what they learn.

[Jordan]: Finally, new reporting on a University of Maryland trial shows why implementation details can matter as much as the tool itself.

[Jordan]: In the randomized trial of Maryland’s Virtual Study Assistant, only about 15 percent of students assigned access used it. Among those users, 73.8 percent asked for information, explanations, or solutions.

[Jordan]: Only 11.3 percent used it for practice or test preparation. And just 0.7 percent sought feedback on their own work.

[Jordan]: The tool’s default mode returned correct answers immediately. Relatively few instructors enabled a tutor mode designed to guide students toward answers. Researchers also noted that visible chat logs, along with access to other AI tools, may have affected student behavior.

[Jordan]: So what should we conclude? Not that AI tutoring simply worked or failed.

[Jordan]: The more useful conclusion is that configuration, instructor preparation, privacy expectations, student incentives, and the design of the learning activity all shaped the result.

[Jordan]: If the goal is guided practice, a default setting that gives the answer may be working against that goal. Access alone isn’t an implementation strategy.

[Jordan]: Across these developments, a consistent picture is emerging. AI may be most valuable when it supports practice, feedback, scaffolding, and reflection. It becomes more problematic when it replaces the first encounter with knowledge, the struggle involved in making sense of an idea, or the independent application of learning.

[Jordan]: For institutions, that means the central work isn’t choosing between enthusiasm and resistance. It’s designing conditions in which AI serves a clear educational purpose—and then checking whether it actually does.

[Jordan]: Here are three ideas worth considering.

[Jordan]: First, design AI pilots around learning processes rather than around tools.

[Jordan]: A pilot proposal could identify the learning process AI is meant to support, the human activity it will augment, the work students must still do independently, and the evidence that would count as success.

[Jordan]: Second, consider making guided practice the default for student-facing tutors.

[Jordan]: That could mean asking students to attempt a problem, explain their reasoning, or respond to a hint before the system reveals a solution. Evaluation should look beyond usage and satisfaction. It should also examine the quality of student reasoning and later independent performance.

[Jordan]: And third, create small, evidence-oriented course redesign funds.

[Jordan]: Stanford’s grants and Iowa’s supported pilots suggest a practical model: modest funding, a bounded intervention, faculty-development support, student feedback, and a short report on learning outcomes, workload, equity, and unintended effects.

[Jordan]: None of these approaches requires an institution to predict the entire future of AI. They do require clarity about what students are supposed to learn, patience with evidence, and a willingness to redesign when the technology changes the learning process in unexpected ways.

[Jordan]: That’s all for this episode of AI in Higher Education. I’m Jordan. Thanks for listening.
[Jordan]: The voice in this episode is AI-generated.

Designing Evidence-Based AI Pilots for Higher-Education Learning
Broadcast by