"PrepMate AI: a local AI interviewer built for my friend"
** What I Built** I built PrepMate AI , a local AI-powered interview practice partner for my college friend who is preparing for software development interviews. While preparing for interviews, he found it difficult to practice realistic mock interviews because he did not always have someone available to ask technical questions and give feedback on his answers. I wanted to build something that…
I developed an AI-powered interview practice tool called PrepMate AI for my college friend who was preparing for software development interviews. Traditional methods of practicing interviews often relied on finding someone available to ask technical questions and provide feedback, which was not always feasible. My goal was to create a solution that addressed this specific challenge rather than developing a general-purpose chatbot.
PrepMate AI is an AI interviewer that conducts technical interviews, evaluates answers, asks follow-up questions, and generates a final performance report. The user experience is straightforward: select a topic, choose the difficulty level, start the interview, and interact with the AI-generated questions and responses. The demo can be accessed at [https://prep-mate-ai-seven.vercel.app](https://prep-mate-ai-seven.vercel.app).
The core of PrepMate AI is a locally running open-weight model, specifically Qwen3 4B, accessed through the Ollama platform. This setup allows the application to function without relying on remote proprietary AI APIs for each question. The interview flow involves choosing the interview category, difficulty, and number of questions; receiving an AI-generated question; submitting an answer; receiving AI evaluation and feedback; and continuing with follow-up questions based on the evaluation.
After the interview concludes, a final performance report is generated, summarizing the candidate's performance and providing question-by-question feedback.
Technically, PrepMate AI is a full-stack application built around a locally running open-weight model. The frontend is constructed with React and Vite, while the backend utilizes Node.js and Express. The AI integration is facilitated by Ollama and the Qwen3 4B model, communicating via a REST API. The architecture comprises the React frontend, Node.js and Express backend, Ollama running on localhost:11434, and the Qwen3 4B model for local inference.
Building PrepMate AI required setting up the interview configuration, generating AI questions, evaluating answers, and handling follow-up questions. The interview setup allows users to choose the interview category, difficulty, and number of questions. When the interview begins, the backend sends a request to Qwen3 to act as a technical interviewer, generating a question based on the selected category, difficulty, interview context, and previous questions and answers.
Following the answer submission, the response and relevant interview context are sent to Qwen3, which evaluates the answer and returns a score, verdict, strengths, issues, ideal answer, and a follow-up question. These details are then displayed as structured feedback within the application.
To enhance realism, PrepMate AI incorporates follow-up questions that mimic a more natural conversation flow, moving beyond the typical question-answer-next question sequence. Once the interview is finished, PrepMate AI compiles a final performance report, which includes an overall performance score, average score, strengths, areas for improvement, recommended topics, and question-by-question feedback.
The decision to run the model locally using Ollama and the Qwen3 4B model offered several advantages for PrepMate AI. It provided complete control over the AI, enabling customization of prompts, interview context, conversation history, and evaluation format. Additionally, running the model locally enhanced privacy, as interview answers containing personal information about the candidate's preparation and progress remained on the user's machine rather than being transmitted to a cloud-based service.
This local inference approach was crucial for the project's development over a weekend, allowing for rapid experimentation without the constraints of paid requests.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.