This paper synthesizes promising practices from a four-team research and development initiative focused on improving the performance of AI systems for K–12 teaching and learning.
Drawing on projects involving AI tutoring, teacher feedback, and benchmark development, the authors identify technical and partnership practices that can increase the usefulness and reliability of AI tools in educational settings.
The paper highlights the importance of carefully designed datasets and evaluation frameworks, with particular emphasis on pedagogical grounding rather than scale alone. It also examines how education researchers and frontier model developers can divide responsibilities and expertise in ways that strengthen both technical performance and educational relevance.
Across the projects, the findings point to the value of combining technical experimentation with domain expertise, clear research questions, and evaluation approaches that reflect the needs and constraints of real teaching and learning environments.
The practices are offered as an evolving resource for researchers and developers as AI model capabilities, evaluation methods, and research infrastructure continue to advance.