AI-Driven Test Grading with Ext JS, Node.js, and OpenAI: A Secure Enterprise Architecture
Get a summary of this article:
Learn how to build an AI-driven test grading system with Ext JS, Node.js, MySQL, and OpenAI using secure API architecture, deterministic scoring, and human oversight.

Automated assessment has become a key capability for modern EdTech platforms, but evaluating open-ended responses differs fundamentally from scoring multiple-choice questions.
A multiple-choice response can be evaluated with deterministic logic. An open-ended answer requires interpretation: Does the student understand the underlying concept? Did they address the required criteria? Should partial credit be awarded? Does the response satisfy the grading rubric without relying on assumptions?
At small scale, educators can manage this complexity manually. At enterprise scale, the model breaks down.
Long grading sessions introduce human fatigue. Evaluation standards can gradually shift during a session, creating what the source architecture describes as grading creep—where later responses may be judged relative to previously reviewed answers rather than consistently against the defined rubric.
The architectural opportunity is therefore not simply to “add AI to grading.” It is to build an AI-driven test grading system in which AI evaluation is surrounded by deterministic application logic, secure API boundaries, structured data contracts, server-side calculations, and human oversight.
This article examines one such architecture using Sencha Ext JS, Node.js, MySQL, and OpenAI.
The Enterprise Challenge: Scaling Open-Ended Assessment
For an EdTech platform, the problem is larger than generating a score.
A production-grade grading system must answer several questions simultaneously:
- How is the student’s answer evaluated against a defined rubric?
- How are partial scores determined?
- How are AI-generated results validated?
- How are API credentials protected?
- Where does business logic reside?
- How are scores persisted?
- How can educators audit or override an AI-generated result?
- How does the interface remain synchronized with backend state?
These requirements make the architecture more important than the model itself.
The proposed solution separates the system into three major application tiers:
Ext JS DataGrid → Node.js middleware → MySQL, with the Node.js service acting as the controlled integration point for the OpenAI API.
This separation keeps presentation, business logic, persistence, and AI execution independently governed.
A Reference Architecture for AI Test Grading
The core architecture can be represented as:
┌──────────────────────────┐
│ Ext JS DataGrid │
│ │
│ Tests │
│ Student Attempts │
│ Answer Details │
│ AI Feedback │
└────────────┬─────────────┘
│
REST / JSON
│
▼
┌──────────────────────────┐
│ Node.js / Express │
│ │
│ Routes │
│ Controllers │
│ Services │
│ Repositories │
└───────┬──────────┬───────┘
│ │
MySQL │ │ OpenAI API
│ │
▼ ▼
┌────────────┐ ┌────────────┐
│ MySQL │ │ OpenAI │
│ Database │ │ API │
└────────────┘ └────────────┘
The architectural principle is straightforward: the browser manages interaction; the backend manages trust.
The Ext JS application provides the desktop-grade interface instructors need to inspect assessments and feedback. Node.js owns the evaluation workflow and external API communication. MySQL provides persistent storage for tests, questions, students, attempts, answers, scores, and feedback.
That separation becomes particularly important when AI is introduced into an application that handles educational records.
Ext JS as the Operational Interface
The frontend is built around the Sencha Ext JS DataGrid, providing instructors with a structured environment for managing assessments and reviewing AI-generated evaluations.
Rather than treating AI grading as an isolated background process, the interface makes the evaluation workflow observable.
The application can expose:
- Student assessment attempts
- Test configurations
- Individual questions
- Student responses
- AI-generated scores
- Feedback
- Strengths and weaknesses
- Evaluation status
- Manual overrides
The architecture uses an Ext.grid.Panel as the primary DataGrid, supported by action columns, double-click interactions, modal detail windows, row expanders, nested grids, and dynamic store reloads.
This creates an important enterprise characteristic: AI does not disappear behind an API call. Its output becomes part of an auditable application workflow.
An instructor can open an assessment attempt, inspect individual responses, review the generated feedback, and determine whether an intervention is necessary.
Node.js as the AI Control Layer
The Node.js Express application forms the middle tier between the browser, database, and AI service.
The backend is deliberately separated into four architectural layers:
Routes
Routes define the available API endpoints and map incoming requests to controllers.
Controllers
Controllers manage request parameters, HTTP status codes, and response formatting.
Services
The service layer contains the business logic, coordinates the grading workflow, constructs AI requests, and manages the OpenAI integration.
Repositories
Repositories isolate database operations and SQL queries from the rest of the application.
This structure prevents the frontend from becoming responsible for business-critical evaluation logic.
More importantly, it creates a clear control boundary around the AI service.
MySQL as the System of Record
AI-generated evaluation should not become the system of record.
The database remains responsible for storing the authoritative assessment state.
The source architecture uses five core tables:
| Table | Purpose |
|---|---|
| tests | Assessment configuration |
| questions | Questions, reference answers, rubrics, and maximum scores |
| students | Enrolled student information |
| test_attempts | Overall assessment sessions and aggregate results |
| student_answers | Individual responses, scores, feedback, strengths, and weaknesses |
The relational model connects attempts to students and tests, while individual answers remain linked to both their assessment attempt and question. This structure ensures that an AI-generated score can be traced back to the specific response and grading criteria that produced it.
That traceability matters.
An enterprise grading system should be able to answer not only “What score did the student receive?”, but also “Which question produced this score, which criteria were applied, and what feedback was generated?”
Security Starts Before the AI Model
One of the most important architectural decisions is also one of the simplest:
Never expose the OpenAI API key to the browser.
Calling an LLM directly from an Ext JS application creates a credential-management problem. Browser-based credentials can potentially be inspected through developer tools, client-side memory, network activity, or improperly secured repositories. The consequences can include unauthorized API usage, exhausted quotas, and unexpected costs.
The safer architecture places Node.js between Ext JS and OpenAI:
Ext JS Browser
│
│ REST / HTTP POST
▼
Node.js Express
│
│ Private API Key
▼
OpenAI API
The backend becomes an AI isolation gateway.
API credentials remain server-side and are loaded through environment configuration rather than being embedded in frontend code.
The architecture also applies additional controls:
- .env protection
- .gitignore safeguards
- Origin restrictions through cors
- Security headers through helmet
- Request validation
- Backend-only AI communication
The source implementation demonstrates this pattern with Express, Helmet, CORS, dotenv, and the OpenAI client.
For enterprise applications, this distinction is fundamental:
The frontend can request an evaluation. It should never own the credentials or authority required to execute that evaluation.
Prompt Engineering for Controlled AI Evaluation
Security protects the API boundary. It does not, by itself, make AI grading reliable.
The second architectural challenge is evaluation consistency.
Large language models are probabilistic systems. An enterprise grading workflow therefore needs a structured prompt contract that constrains how the model interprets questions, evaluates responses, and returns results.
The source architecture separates the prompt into five components:
1. Role
The model receives an explicit role, such as an expert, strict, and fair technical instructor.
2. Rules
The evaluation rules define how answers must be judged.
The model is instructed to evaluate each response independently, avoid unstated assumptions, and follow the supplied rubric.
3. Payload
The model receives structured assessment data:
- Question
- Student answer
- Reference answer
- Grading criteria
- Maximum score
4. Schema
The response format is explicitly defined so that Node.js can parse the result programmatically.
5. Limits
The prompt establishes score boundaries and prevents the model from inventing additional criteria or exceeding the defined maximum score.
This is a critical shift in mindset.
Prompt engineering in an enterprise application is not merely about improving AI responses. It is about defining an operational contract between the model and the application.
Designing for Consistent Scoring
A grading system should not reward an answer simply because it sounds convincing.
The source architecture establishes two important scoring controls.
No unstated assumptions
If the grading rubric requires a particular concept and the student does not explicitly demonstrate it, the system should not assume that the student knew it.
Controlled partial credit
Partial credit should correspond to explicitly defined rubric elements rather than subjective interpretation.
This approach turns the rubric into the primary evaluation authority, with the AI acting as the mechanism for interpreting the student’s response against that rubric.
The distinction is important:
AI interprets. The rubric constrains. The application validates.
Structured JSON Creates an Application Contract
Free-form AI responses are difficult to integrate reliably into production systems.
A structured response contract creates a much stronger boundary.
The proposed evaluation response contains fields such as:
{
"attemptId": 8,
"testId": 2,
"passed": false,
"totalScore": 48,
"maxScore": 80,
"percentage": 60,
"evaluations": [
{
"questionId": 101,
"answerId": 405,
"score": 6,
"maxScore": 10,
"feedback": "...",
"strengths": "...",
"weaknesses": "..."
}
]
}
The important design principle is that the response is machine-readable.
Node.js can parse the result, validate it, recalculate the aggregate metrics, and persist the approved values into MySQL.
Never Let the LLM Own the Mathematics
One of the strongest architectural decisions in the system is to separate qualitative evaluation from quantitative calculation.
The model evaluates individual responses.
Node.js calculates the final score.
The backend iterates over the returned evaluations, calculates the total score and maximum score, derives the percentage, and determines whether the student passed according to the assessment’s configured threshold.
Conceptually:
AI
│
├── Question 1 → 8 / 10
├── Question 2 → 6 / 10
├── Question 3 → 9 / 10
│
▼
Node.js
│
├── Calculate total
├── Calculate percentage
├── Apply passing threshold
└── Persist verified result
This creates a clear division of responsibility.
The AI provides assessment judgments. The application performs deterministic arithmetic.
That principle extends beyond education. Whenever an LLM participates in a workflow containing financial, mathematical, compliance, or threshold-based decisions, deterministic application logic should own the final calculation wherever possible.
The End-to-End AI Grading Workflow
The complete grading workflow follows a controlled sequence:
Instructor triggers evaluation
↓
Express route receives request
↓
Attempt controller validates request
↓
Repository retrieves attempt + answers
↓
Service constructs structured payload
↓
OpenAI evaluates responses
↓
Node.js parses structured JSON
↓
Node.js recalculates totals
↓
Validated results saved to MySQL
↓
Ext JS stores reload
↓
Instructor reviews results
The source implementation uses an endpoint such as:
POST /attempts/:id/grade
The service gathers the relevant assessment information, constructs the evaluation payload, invokes the OpenAI completion endpoint, parses the structured response, performs the mathematical safeguard, and commits the resulting scores and feedback through the repository layer.
This makes the workflow observable and testable at every stage.
Keeping the Ext JS Interface in Sync
AI grading introduces another engineering challenge: state synchronization.
The instructor may trigger an evaluation from inside a nested details window while the parent DataGrid remains visible.
A full page refresh would be disruptive and unnecessary.
The Ext JS architecture instead reloads the relevant stores after the backend completes the evaluation.
The workflow is:
- Instructor opens an assessment attempt.
- A DetailsWindow displays individual answers.
- The nested answer grid loads the attempt’s responses.
- Instructor triggers evaluation.
- Ext.Ajax.request sends the request to the backend.
- Node.js performs the grading workflow.
- The nested answer store reloads.
- The parent assessment grid reloads.
- Updated scores and feedback become visible.
The source implementation specifically uses the view controller to coordinate these nested component updates and clears the loading state across both success and failure paths.
This is where a rich enterprise JavaScript framework becomes valuable: AI processing can be integrated into a stateful operational interface rather than bolted onto an otherwise static application.
Human-in-the-Loop Is a Governance Feature
AI-driven grading should not be designed around the assumption that every model output is automatically correct.
The architecture therefore retains an explicit human-in-the-loop workflow.
If an evaluation response is malformed, backend exception handling prevents invalid data from being committed.
If an AI-generated score requires review, an educator can inspect the response, examine the feedback, and override the stored score.
This creates a three-stage governance model:
AI Evaluation
↓
Deterministic Validation
↓
Human Review / Override
The goal is not to remove educators from the process.
The goal is to move educators away from repetitive first-pass evaluation and toward exception handling, quality assurance, and academic judgment.
That distinction is strategically important for enterprise EdTech platforms.
Managing AI Costs at Scale
AI grading also introduces an operational economics question.
Every evaluated response consumes model resources. At scale, token consumption becomes an architecture concern rather than simply an API expense.
The source identifies two execution strategies.
Granular evaluation
Each student’s assessment is submitted as an individual evaluation request.
Advantages include:
- Lower per-request context size
- Faster individual processing
- Real-time evaluation workflows
- Simpler failure handling
Batch evaluation
Multiple student attempts are grouped into a larger request.
This can reduce repeated prompt and request overhead, but increases context size and processing time. It also requires careful management of model context limits.
The right approach therefore depends on the operational objective.
For interactive instructor workflows, granular evaluation may align better with responsiveness. For large-scale asynchronous grading, batching may offer a different cost-performance profile.
Configuration Consistency Matters
Production AI systems need controlled configuration.
The source architecture uses:
temperature: 0
to minimize output variability and also identifies consistent seed configuration as a mechanism for improving repeatability when the same input is evaluated multiple times.
The broader architectural principle is more important than any individual parameter:
AI evaluation should be treated as a controlled production workload, not an unrestricted conversational interaction.
That means versioning and governing:
- System prompts
- Evaluation rubrics
- Output schemas
- Model configuration
- Validation rules
- Error-handling behavior
- Human override processes
Building a Production-Ready AI Grading Platform
The technical components are individually familiar.
Ext JS provides the application interface.
Node.js provides the backend orchestration.
MySQL provides persistence.
OpenAI provides language-model evaluation.
The value comes from how these components are combined.
A production-ready architecture establishes clear ownership:
| Responsibility | System Layer |
|---|---|
| Instructor interaction | Ext JS |
| Assessment visualization | Ext JS DataGrid |
| API routing | Node.js |
| Business logic | Node.js Services |
| AI orchestration | Node.js Services |
| Deterministic calculations | Node.js |
| Persistent assessment data | MySQL |
| Qualitative evaluation | OpenAI |
| Final review | Educator / HITL |
This division creates a system in which no single component is responsible for everything.
What Enterprise EdTech Leaders Should Take Away
The most important lesson from AI-powered assessment is that the model is only one component of the system.
An enterprise-grade implementation requires architecture around the model.
That architecture should provide:
Secure AI integration
AI credentials remain behind a controlled backend boundary.
Structured evaluation
Prompts define explicit roles, rules, inputs, schemas, and scoring constraints.
Deterministic application logic
The backend—not the model—calculates aggregate scores and pass/fail outcomes.
Traceable persistence
Every evaluation remains associated with a specific assessment, question, student response, and result.
Rich operational interfaces
Educators can inspect AI-generated feedback directly within the Ext JS DataGrid workflow.
Human oversight
AI-generated results remain reviewable and overridable.
Operational governance
Token consumption, configuration, errors, retries, and evaluation consistency can be managed as production concerns.
These principles transform AI grading from a simple API integration into an enterprise application architecture.
Conclusion
The future of automated assessment is not simply about asking an AI model to grade more student responses.
It is about designing systems where AI can operate at scale without becoming the uncontrolled source of truth.
A combination of Sencha Ext JS, Node.js, MySQL, and OpenAI provides a practical architecture for achieving that separation.
Ext JS provides the interactive React Data Grid and instructor experience. Node.js creates the secure orchestration layer and owns deterministic business logic. MySQL maintains the authoritative assessment record. OpenAI performs structured qualitative evaluation. Human-in-the-loop controls provide the final governance layer.
The result is a system designed around a simple enterprise principle:
Automate the evaluation workload. Keep the authority, validation, and governance in the application.
For EdTech platforms, that distinction can make the difference between an AI experiment and a production-ready assessment capability.
Enterprise analytics is entering a new phase. Business users increasingly expect applications to answer questions…
Sencha’s AI assistant has grown from a focused Ext JS helper into a product-aware chatbot…
Enterprise applications rarely fail because a grid cannot render a few hundred records. The real…




