Photo by Bruno Kelzer on Unsplash

What Districts Really Look For When Buying AI Tools

(and What They Wish They Could Look For)

Zia Hassan (JHU)

Definitions

K-12 AI Tools

Educational technology tools that market themselves as AI-enabled

Efficacy Evidence

Data or findings from a study showing whether a tool improves student learning outcomes

Study Overview

  • K–12 districts are adopting AI-enabled EdTech faster than researchers can establish evidence of efficacy
  • This study examines what signals districts rely on when procuring AI tools, and how much weight efficacy evidence carries relative to those signals

Methods

  • 5 semi-structured interviews across different K–12 districts, paired with a review of the literature on AI-enabled EdTech procurement and evidence of efficacy
  • Thematic coding of interviews
  • Analysis of interviews to support or contradict literature

The Headline Finding

  • Districts lean on practical, available signals while evidence of AI tools' efficacy catches up
  • Two signals act as thresholds a tool must clear before efficacy is even discussed
  • Two more act as proxies, standing in for efficacy evidence that isn't readily available
  • All five interviewees said efficacy evidence matters — but none raised it first

Threshold Signals

  • Functional validation – "does it work as intended?"
  • Interoperability – "does it connect to our systems?"
  • A tool that fails either test is disqualified, regardless of any other evidence

"We'll start playing with importing information in, seeing what it looks like on the student side… and putting in things to see what the generated result's going to be."

"It had to go through Canvas… to make sure everything would be matching."

Proxy Signals

  • Peer adoption – "what are similar, trusted districts using?"
  • User delight – "was it enjoyable to use?"
  • Both are within a district's control and immediately available — unlike efficacy evidence

"A salesman's gonna tell you what you wanna hear... to sell the product. But if you hear it from a source that says, hey, this is really good... you can trust that as a valid source."

"If the students like it, that's good enough for our teachers, and our teachers will advocate on their behalf."

Why the Gap Exists: Evidence Is Still Scarce

  • Of 800+ papers on AI in K–12, only 20 met the bar for strong causal evidence (Fesler et al., 2026)
  • None of the student-facing causal studies took place in U.S. K–12 schools; most research is postsecondary; effect is mitigated when compared with non-AI interventions (Létourneau et al. 2025)
  • No research yet examines impacts for students with IEPs or 504 plans

What Evidence Do We Have?

  • With teacher support during students' GenAI use: large effect (g = 1.43). Without it: essentially none (g = 0.08) (Gu et al., 2025)
  • But how does AI + teacher compare to the same teacher support without AI? We don't know yet.

What Evidence Do We Have?

  • After correcting for publication bias, the overall effect of GenAI on STEM learning drops to near zero (Boolzen et al., 2026)
  • 33 of 59 comparisons pitted an interactive AI activity against a passive control (e.g., watching a lecture video)

Why the Gap Exists: Vendor Evidence Isn't Trusted (Morrison et al., 2014)

  • Fewer than half of district respondents in prior research trusted vendor-submitted evidence
  • Interviewees described vendor claims as unverifiable on their face
  • Functions more as institutional cover than as a real decision driver

"Yeah, your Sharpie can do that… show me the real stuff."

Why the Gap Exists: Peer Choices & Limited Frameworks

  • Districts substitute peer behavior for their own efficacy evaluation, AKA institutional isomorphism (DiMaggio & Powell, 1983)
  • Existing frameworks (EdSafe, ISTE, CoSN) are used inconsistently
  • AI tools are adopted before they can really be evaluated

How the Signals Connect

Recommendations to move toward evidence of efficacy

  • A shared, pre-vetted evidence repository
  • Replication studies in underrepresented districts
  • More studies that compare non-AI teacher interventions to AI interventions with teacher supervision
  • A standardized, rapid-pilot protocol

Recommendations for threshold markers (based on interview data)

  • Developing a common certification effort for interoperability (works with Canvas/Google/Blackboard)
  • Capacity-building for small & rural districts. Seed shared regional evaluation capacity.
  • An online forum or place for exchange about what has worked at various schools

Appendix Matrix: Impact vs. District Burden

Recommendation Impact District Cost/Burden Cultural Shift Time to Benefit
Shared evidence repository Every district gets a tiered, efficacy-verified tool list Low–Medium Moderate — trusting an external rating over peer/vendor signals Slow — value accumulates over time
Replication studies Tests whether findings generalize to underrepresented populations Moderate–High Low–Moderate Slow — 1–3 years to design, run, publish
Standardized rapid-pilot protocol Streamlines something districts already do inconsistently Low–Moderate Low — doesn't change what districts do, just how rigorously they document it Fast — no new infrastructure required first

Appendix Matrix: Evidence Basis

Recommendation Key Evidence
Shared evidence repository Peer adoption is already the dominant signal (DiMaggio & Powell, 1983); Digital Promise's League of Innovative Schools shows districts adopt ESSA-tier certifications
Replication studies K-12 meta-analyses (Yi et al., 2025; Wu et al., 2026) code grade level, sample size, design — but not SES, Title I, or IEP/disability status
Standardized rapid-pilot protocol Districts pilot inconsistently, nearly half view teacher-led pilots as a "nuisance" (Adams-Bass et al., 2015); small/rural districts lack dedicated evaluation staff (Bacak et al., 2023)

Contact

zhassan4@jh.edu

linkedin.com/s/zia-s-hassan

Slides available at ziahassan.com/sree2026

References

Adams-Bass, V., Atchison, D., & Moore, L. (2015). Pilot-to-purchase project. UC Davis School of Education.

Bacak, J., Wagner, J., Martin, F., Byker, E., Wang, W., & Ahlgrim-Delzell, L. (2023). Examining technologies used in K-12 school districts: A proposed framework for classifying educational technologies. Journal of Educational Technology Systems. https://doi.org/10.1177/00472395231155605

Boolzen, C., Kuhn, J., Flegr, S., Rott, E.-M., Stausberg, N., & Küchemann, S. (2026). Evidence of impact and interpretational limits of generative AI in STEM education: A systematic review and meta-analysis on cognitive learning outcomes. Artificial Intelligence Review.

DiMaggio, P. J., & Powell, W. W. (1983). The iron cage revisited: Institutional isomorphism and collective rationality in organizational fields. American Sociological Review, 48(2), 147.

Fesler, L., Martinez Claeys, J. P., Agnew, C., & Loeb, S. (2026). The evidence base on AI in K-12: A 2026 review. Stanford SCALE Initiative.

Gu, J., & Yan, Z. (2025). Effects of GenAI interventions on student academic performance: A meta-analysis. Journal of Educational Computing Research, 63(6), 1460–1492.

Létourneau, A., Deslandes Martineau, M., Charland, P., Karran, J. A., Boasen, J., & Léger, P. M. (2025). A systematic review of AI-driven intelligent tutoring systems (ITS) in K-12 education. npj Science of Learning, 10, 29.

Lue, G. (2026, January 16). Blueprints for better AI in education: Insights from the K-12 infrastructure program. Digital Promise.

Morrison, J. R., Ross, S. M., & Corcoran, R. P. (2014). Fostering market efficiency in K-12 ed-tech procurement.

Wu, X., Zhu, P., Zhang, J., Yin, M., & Wang, Y. (2026). ChatGPT's impact on student learning outcomes: A meta-analysis of 35 experimental studies. Humanities and Social Sciences Communications, 13(1), 684.

Yi, L., Liu, D., Jiang, T., & Xian, Y. (2025). The effectiveness of AI on K-12 students' mathematics learning: A systematic review and meta-analysis. International Journal of Science and Mathematics Education, 23(4), 1105–1126.