AI Response Evaluation
Assessing model outputs for accuracy, relevance, clarity, factual consistency, and instruction following using structured project guidelines.
AI Evaluation Specialist · Prompt Engineer · Web Developer · Penetration Tester
I help evaluate AI systems, refine prompts, build structured digital workflows, and create professional web experiences with a focus on clarity, evidence, security, and long-term usability.
In Focus
Seven perspectives on the disciplines, places, and mindset behind the work.

Structured, rubric-based review of model outputs.
By the Numbers
0+
YouTube Subscribers
Grown through AI-assisted storytelling and video production.
0+
Total Views
Across AI-driven content channels built from scratch.
0+
Years in AI Evaluation
Across Remotasks and TELUS Digital, remote.
0
Core Disciplines
Evaluation, prompting, training, web, security, documentation.
About
How evidence-based thinking connects evaluation, prompt design, development, and cybersecurity into one coherent practice.
Obasiochie Vincent Chimaobi is an experienced technology expert with a background spanning AI response evaluation, prompt engineering, data annotation, generative AI workflows, web development, and cybersecurity-informed analysis.
His evaluation mindset treats every project guideline as a judgement recipe — breaking instructions into clear requirements before rating outputs. He analyses each criterion by identifying what needs to be checked, what counts as full or partial fulfilment, and what evidence is required, then applies a structured five-step rubric approach to reach a final rating.
That same discipline carries into prompt engineering, where objectives become precise instructions with defined context, constraints, and output formats — and into web development, where responsive, accessible, and maintainable structure matters as much as visual polish. Cybersecurity work adds a risk-aware lens to everything: careful verification, error detection, and disciplined documentation.
Professional Strengths
B.Sc. Physics Electronics
University of Port Harcourt
Rivers State, Nigeria
Graduated: 2021
Udemy Prompt Engineering Certification
Junior Penetration Tester (eJPT)
English
Igbo
Belfast, Northern Ireland, United Kingdom
Independent Client Engagements / HackerOne / Bugcrowd
Remote
Independent AI Content Projects
Remote
TELUS Digital
Remote
Remotasks
Remote
The Journey
A career built on analytical thinking, evidence-based judgement, and continuous learning — one milestone at a time.
University of Port Harcourt
Began a rigorous programme combining physics, electronics, and computational thinking — the foundation for a technical career.
Remotasks
Entered the AI industry through data annotation and quality control — learning how training data shapes model behaviour.
University of Port Harcourt
Graduated with B.Sc. Physics Electronics. Final-year project: an underwater laser detection security system using Raspberry Pi and Python.
TELUS Digital
Promoted to evaluating AI responses with structured rubrics — factuality, hallucination detection, instruction following, and comparative judgement.
Independent Projects
Built AI-driven YouTube channels using prompt engineering and generative tools — reaching 29,000+ subscribers and 2M+ views.
HackerOne & Bugcrowd
Began authorised penetration testing and responsible disclosure — bringing security-informed thinking into all technical work.
Expertise
Twelve disciplines that work together — from AI evaluation and prompt engineering to web development and cybersecurity-aware thinking.
Assessing model outputs for accuracy, relevance, clarity, factual consistency, and instruction following using structured project guidelines.
Converting business objectives into clear, constraint-aware instructions that improve reliability and reduce ambiguity in AI workflows.
Supporting machine-learning training workflows by preparing, labelling, and validating human-reviewed data with consistent quality.
Human-in-the-loop labelling and review that keeps model behaviour aligned with project-specific accuracy standards and guidelines.
Designing and building responsive, accessible websites with a focus on information architecture, performance, and maintainable structure.
Authorised security research, responsible disclosure, and risk-aware development thinking brought into every technical decision.
Clear, evidence-based written communication that makes complex systems, findings, and processes easy to understand and act upon.
Building end-to-end generative pipelines across text, image, and video for research, content, and storytelling use cases.
Breaking problems into measurable parts, forming evidence-based judgements, and refining conclusions as new data arrives.
Identifying unsupported claims, fabricated facts, and weak reasoning in model outputs before they reach end users.
Designing and applying structured rubrics that turn subjective quality into repeatable, comparable, and auditable judgements.
Combining research, disciplined testing, and structured reasoning to resolve ambiguous technical challenges end to end.
How I Work
A personal evaluation methodology developed over thousands of rated responses — structured, repeatable, and auditable.
Before rating anything, I form a clear picture of what a fully correct response should look like — every constraint, format, and edge case considered.
I read the response in full and form an instinctive impression. This initial read is useful, but never the final word — evidence must confirm it.
I break the guideline into discrete criteria, identifying what must be checked, what counts as full or partial fulfilment, and what evidence is required.
I map the response back against each criterion, citing specific passages. Unsupported claims and hallucinations surface at this stage.
I assign the rating based on guideline alignment, document the reasoning, and flag anything that needs a second review for consistency.
This approach treats every guideline as a judgement recipe — breaking instructions into clear requirements before rating, so each decision is defensible and consistent.
Proficiency
A transparent view of where the deepest expertise lies — from AI evaluation and prompt engineering to web development and security research.
Percentages reflect depth of hands-on experience, not formal certification scores. Updated as new skills develop.
Toolkit in Orbit
A constellation of the platforms, languages, and security tools used across evaluation, prompting, development, and research.
ChatGPT, Claude, Gemini, Midjourney, Veo 3.1, and Nano Banana — used daily for response evaluation, prompt development, content generation, and AI-assisted video workflows.
Python, Next.js, React, and Raspberry Pi — covering web development, automation, and sensor-based security systems built from the ground up.
HackerOne and Bugcrowd for responsible disclosure, with evidence-led methodology and remediation-focused documentation.
Hover the constellation to pause the orbit. Each tool is positioned by discipline — evaluation at the centre, generative platforms in the inner ring, development and security in the outer ring.
Full Stack
A transparent view of the full technology stack — filter by category and see exactly where the expertise lies.
Daily use for evaluation, prompting, and content workflows.
Complex reasoning tasks and long-context evaluation.
Multimodal evaluation and research assistance.
Generative image creation for content workflows.
AI video generation for storytelling content.
Production React framework with App Router.
Component architecture and state management.
Type-safe development across the stack.
Utility-first styling for rapid, consistent UI.
Semantic, accessible markup foundations.
Server-side JavaScript for APIs and tooling.
Scripting, automation, and security tooling.
Type-safe database access and migrations.
Lightweight databases for static-friendly apps.
Responsible vulnerability disclosure platform.
Crowdsourced security testing programmes.
Hardware security projects and sensor systems.
Version control and collaborative workflows.
CMS for client websites and content management.
Principles
Six principles that shape every decision — from how an AI response is rated to how a vulnerability is documented.
Every judgement — an AI rating, a security finding, a design decision — should be backed by evidence you can point to. Opinions are starting points, not conclusions.
Show the evidence, then the conclusion.
Clear writing, clear structure, clear reasoning. Ambiguity wastes time and breeds mistakes. Whether in a prompt, a report, or a line of code — be understood the first time.
If it needs reading twice, rewrite it.
A risk-aware mindset belongs in every project, not just penetration tests. Small decisions compound into real safety — or real vulnerability.
Threat-model first, build second.
Code, prompts, and findings all outlive the moment they're created. Document the reasoning so the next person — or future you — understands not just what, but why.
Future you is a colleague. Be kind.
Not every problem needs a framework. Not every site needs a dashboard. Match the solution to the actual need — simple stays simple, complex earns its complexity.
The best tool is the one that fits.
Showing up daily — evaluating, refining, documenting, learning — matters more than bursts of intensity. Small, steady work builds real expertise over time.
Discipline beats motivation.
AI Evaluation
Structured, rubric-based review of model outputs: accuracy, factuality, instruction following, and hallucination risk assessed with care.
Checking whether a response actually does what the instruction asked — every constraint, format, and edge case accounted for.
Verifying claims against trusted sources and flagging statements that are unsupported, outdated, or internally inconsistent.
Spotting fabricated facts, invented citations, and confident-but-wrong reasoning before they reach an end user.
Applying structured rubrics so quality is measured consistently, compared fairly, and audited later.
Reviewing multiple candidate outputs side by side and selecting the strongest based on evidence, not preference.
Bringing careful, evidence-based human review into automated pipelines to keep model behaviour trustworthy.
Live Demo
A real example of the kind of issues I detect in model outputs — factuality errors, hallucinations, and instruction violations, with the corrected response beside it.
PUser Prompt
Explain the difference between TCP and UDP and give one real-world example of each.
Model Response (Unevaluated)
5 issues foundDetected Issues (5)
“TCP is faster than UDP”Factually incorrect — TCP is slower due to handshake and acknowledgement overhead.
“doesn't check for errors”Violates accuracy — TCP does check for errors via checksums and retransmission.
“invented by Google in 2015”Hallucinated fact — TCP was developed in the 1970s by Vint Cerf and Bob Kahn, not by Google.
“Universal Data Protocol”Hallucinated expansion — UDP stands for User Datagram Protocol.
“Layer 3 of the OSI model”Incorrect — both operate at Layer 4 (Transport), not Layer 3 (Network).
Hover any highlighted snippet or issue card to see the connection. This is the kind of evidence-based, rubric-driven review applied to every AI output I evaluate.
Try the Rubric
This is the actual rubric used to evaluate AI responses. Click score levels for each criterion and watch the weighted total update in real time.
Did the response address every part of the instruction?
Are all claims accurate and supported?
Any fabricated facts or citations?
Is the response clear and well-organised?
Does it fully answer the question?
Excellent
Breakdown
Each criterion is weighted by importance. The final score reflects the weighted average — exactly how production AI evaluation works.
Evaluation Gallery
Real examples of the errors that slip into model outputs — and how evidence-based evaluation catches them. Click any card to see the issue and the fix.
Model Output
According to Smith et al. (2023), TCP was developed at Bell Labs in the 1980s.
The Issue
The cited paper does not exist. The model invented both the author and the publication.
How It's Caught
Verify every citation against a trusted source. Flag any reference that cannot be independently confirmed.
Model Output
Python was first released in 1995 by Guido van Rossum.
The Issue
Python was released in 1991, not 1995. The model stated an incorrect date with full confidence.
How It's Caught
Cross-check factual claims (dates, names, figures) against authoritative sources before accepting them.
Model Output
Here is a 500-word summary of the article: [proceeds to write 500 words]
The Issue
The instruction asked for a summary under 200 words. The model ignored the length constraint entirely.
How It's Caught
Check every constraint in the instruction against the output. Length, format, tone, and scope must all be verified.
Model Output
UDP is faster than TCP because UDP is faster. TCP is slower because it is slower than UDP.
The Issue
The explanation is circular — it restates the claim as the reason without providing any actual explanation.
How It's Caught
Identify circular logic by checking whether the 'reason' adds new information beyond the claim itself.
Model Output
Both TCP and UDP operate at Layer 3 (Network) of the OSI model.
The Issue
Both protocols operate at Layer 4 (Transport), not Layer 3. A fundamental networking error.
How It's Caught
Verify technical classifications against reference documentation — OSI layer assignments are well-defined.
Model Output
TCP is used for video streaming because it is faster, while UDP is used for file downloads because it is reliable.
The Issue
The entire comparison is reversed. TCP is reliable (for files), UDP is fast (for streaming).
How It's Caught
When a response describes trade-offs, verify that the characteristics match the correct entities.
These are illustrative examples of the kind of errors caught during evaluation. Each represents a real category of model failure that rubric-based review is designed to detect.
Prompt Engineering
A practical discipline of defining context, constraints, and output formats, then refining prompts based on observed output failures.
Translating a business goal into a clear, testable instruction with defined context, constraints, and output format.
Specifying exactly what the model should know, what it must avoid, and the shape its answer should take.
Tightening language and structure so outputs become more predictable and less prone to drift.
Refining prompts based on observed output failures — turning each failure into a sharper, more robust instruction.
Building reusable prompt components that support research, content, evaluation, and automation tasks.
Applying prompt discipline across web, research, content, and evaluation work for consistent quality.
Interactive
Compose a structured prompt by selecting a role, task, context, constraint, and format. This is the same approach I use to build reliable, constraint-aware AI workflows.
You are a domain expert with deep technical knowledge. Explain the following concept in detail with examples. The audience is a general, non-technical reader. Support every claim with evidence or examples. Format the output as a bullet-point list.
This is how structured, constraint-aware prompts are built — turning objectives into reliable instructions.
Need a custom prompt framework? Let's talkWeb Development
A growing but credible technical direction — from lightweight static sites to dashboards, support systems, and client portals.
Not every website needs every advanced feature. Simple sites stay lightweight, static, affordable, and professional — while larger projects can grow into dashboards, content managers, or client portals.
Responsive layouts, accessible components, and clean information architecture built with React, Next.js, TypeScript, and Tailwind CSS for speed and maintainability.
Static-friendly architecture that migrates cleanly to professional hosting when a project needs APIs, authentication, or a data layer.
Not every website requires every advanced feature. Simple projects can remain lightweight and affordable, while larger ones grow into dashboards, support systems, or client portals.
Cybersecurity
Authorised security research, responsible disclosure, and risk-aware development, documented with remediation in mind.
Conducting assessments only within agreed scopes, with clear permissions and responsible engagement.
Reporting vulnerabilities through proper channels so they can be fixed before they are exposed.
Capturing clear, reproducible evidence that demonstrates impact without revealing unnecessary detail.
Assessing real-world risk — what an issue enables, who it affects, and how severe it would be if exploited.
Writing findings that lead directly to fixes — clear steps, priorities, and verification guidance.
Bringing a security mindset into every build decision so safety is designed in, not bolted on.
Security Posture
A snapshot of security work — disclosure rates, response times, and evidence quality. All work is done under proper authorisation.
No active threats. Routine monitoring.
Disclosed vulnerabilities with confirmed remediation.
Average time to first response on security reports.
Vulnerabilities disclosed through proper channels.
Average evidence completeness rating on reports.
Metrics are illustrative aggregates. Specific vulnerability details remain confidential under responsible disclosure timelines and NDAs.
Projects
Conceptual project previews representing the kind of work delivered across AI evaluation, prompt engineering, web development, and penetration testing. Click any card for full details and demonstrations.
A structured workflow for evaluating AI responses against rubrics, capturing evidence, and producing comparable quality scores across model versions.
A repeatable framework for turning objectives into constraint-aware prompts, refining them based on observed output failures, and documenting what works.
A responsive, accessible, maintainable personal portfolio built with a modern frontend stack — designed for static deployment and future migration.
A planning framework that helps small businesses choose the right scope — from lightweight static sites to dashboards, support systems, and client portals.
A documentation pattern for responsible vulnerability reporting — evidence, impact, reproduction steps, and remediation guidance without exposing sensitive details.
An end-to-end workflow for producing AI-assisted video and storytelling content — from research and prompt development through review and publication.
Project visuals are conceptual demonstrations representing each expertise area. Click any card to view the full project breakdown and available animation.
Real Impact
Not vanity metrics — each figure represents genuine work, real growth, and measurable outcomes. Click any card to see the detail.
Grown organically through AI-assisted storytelling
Built a channel from zero using prompt engineering and generative AI video workflows. No paid promotion — all organic growth through content quality.
Across AI-driven content channels
Multiple channels reaching audiences interested in AI-generated films, documentaries, and reenactment content.
Evaluation & annotation experience
Continuous work across Remotasks and TELUS Digital, evaluating thousands of AI responses against structured rubrics.
Vulnerabilities reported through proper channels
Security findings documented with evidence, impact analysis, and remediation guidance — submitted via HackerOne and Bugcrowd.
The Difference
A clear comparison of what makes this practice different — not just in skills, but in how the work is done.
Slow responses, vague updates, missing context.
Fast WhatsApp replies, written progress updates, clear scope documents.
Gut-feel ratings, no rubric, inconsistent quality.
Evidence-based rubric scoring, documented reasoning, auditable ratings.
Security as an afterthought, bolted on late.
Risk-aware from day one, threat-modelled approach, responsible disclosure mindset.
Code with no comments, reports with no evidence.
Everything documented — the what, the why, and the evidence behind it.
Vague quotes that creep upward, surprise costs.
Fixed quotes with clear deliverables, no surprises, right-sized scope.
Radio silence after handover, paid support for everything.
30 days included support, questions answered, small tweaks at no extra cost.
Not just deliverables — a fundamentally different way of working. Evidence-based, security-aware, consistently documented, and genuinely supportive after the work is done.
Activity
A year of professional engagement across AI evaluation, prompt engineering, web development, and security research — showing up regularly matters.
Illustrative activity pattern showing consistent engagement across evaluation, development, and security work. Each square represents a day; darker squares indicate more professional activity.
A Day in the Life
A look at the rhythm of a day — balancing evaluation, development, security, and content work with discipline and focus.
Review overnight evaluation queues, prioritise tasks, and check security advisories from disclosed programmes.
Focused rubric-based evaluation of model responses — factuality checks, hallucination detection, and instruction-following review.
Refine prompt templates based on observed output failures, document patterns, and build reusable workflow components.
Break for lunch while reading research papers or documentation on LLM evaluation and security topics.
Build and review responsive, accessible websites — frontend work, component architecture, and performance optimisation.
Authorised penetration testing, vulnerability documentation, and responsible disclosure work on HackerOne and Bugcrowd.
Manage AI-driven content channels, review video outputs, and document the day's findings before signing off.
A flexible framework, not a rigid schedule — the exact rhythm shifts with project deadlines and client priorities, but the balance of evaluation, building, and security work stays consistent.
Website Concepts
Professional website concepts that can be adapted to your brand, colours, content, and business requirements. These are starting points, not completed client projects.
A professional website for a local or small business, with essential pages, contact information, and a clean design that builds trust with customers.
Customizable
A modern, energetic website for a start-up, designed to communicate vision, attract early customers, and showcase the product or service.
Customizable
A website for a service-based business, with clear service descriptions, booking or enquiry options, and trust-building testimonials or case studies.
Customizable
A professional corporate website with multiple departments, team profiles, company information, and structured navigation for larger organisations.
Customizable
A personal portfolio website for a professional, freelancer, or creative, showcasing skills, experience, projects, and contact information.
Customizable
A focused single-page website designed to convert visitors into leads or customers, with a clear call to action and minimal distractions.
Customizable
A concept for an online store, with product listings, shopping cart, checkout flow, and product search. Available as part of the Advanced package scope.
Customizable
A website for a restaurant, cafe, or food business, with menu display, opening hours, location map, and online reservation or ordering options.
Customizable
See a concept you like? It can be customized to match your brand, colours, and business needs.
Tell Me About Your WebsiteConcept designs are professional starting points, not completed client projects. Each concept is adapted to the client's specific requirements before development begins.
Website Packages
Three cumulative tiers for your website project. Each includes everything in the tier below, plus more. Click any package for full details and feature explanations.
A professional entry-level website for customers who need a relatively simple online presence.
9 features included
A more complete website for growing businesses that need stronger content, visibility, and business integrations.
10 features included
A more sophisticated website for customers who need broader functionality, stronger integrations, or advanced capabilities.
10 features included
No active promotion right now. Join the priority list to get notified about the next website offer.
Prices are fixed public prices. Every project is quoted individually based on requirements. Third-party services, hosting, or subscriptions may incur separate costs, which are always clearly disclosed.
Estimator
Select the features you need and get an instant indicative estimate. Every project is quoted individually, but this gives you a realistic starting point.
Estimated range
Indicative estimate only. Final pricing depends on design complexity, content, timeline, and specific requirements. Cybersecurity work is scoped separately.
Credentials
A documented track record across AI evaluation, development, security research, and academic foundations.
Comprehensive training on structured prompt design, context engineering, and iterative refinement for production AI workflows.
Final-year project: underwater laser detection security system using Raspberry Pi and Python. Strong foundation in electronics and programming.
Evaluated AI-generated responses using structured rubrics, assessing accuracy, factuality, instruction following, and hallucination risk.
Contributed to AI training workflows through data labelling, quality control, and rubric-based output evaluation.
Active participant in responsible vulnerability disclosure programmes. Publicly disclosable work includes Syfe.com and University of Port Harcourt.
Built AI-driven YouTube channels reaching 29,000+ subscribers and 2M+ views using prompt engineering and generative AI workflows.
All credentials can be independently verified. Academic transcripts and certification IDs are available on request to serious enquiries.
Always Learning
Technology moves fast. Here's what I'm actively studying to keep my skills sharp and current. Click any card to read the full learning breakdown.
Exploring multi-turn evaluation, adversarial prompting, and red-teaming methodologies for more robust model assessment.
Deepening knowledge of modern web vulnerabilities, secure coding patterns, and automated security scanning workflows.
Mastering the latest React Server Components patterns, streaming, and partial prerendering for production apps.
Writing
Illustrative article previews reflecting the themes I write about as a freelance content writer. Real articles will replace these placeholders.
A practical case for turning subjective AI quality judgements into structured, repeatable rubrics — and the five-step mindset that makes it work.
How treating prompts like living documentation — with context, constraints, and clear output formats — produces more reliable AI workflows.
Why a risk-aware mindset belongs in every project, not just penetration tests — and how responsible disclosure makes the web safer for all.
These are illustrative article previews. Published writing will be linked here as it becomes available.
FAQ
Practical answers to the questions collaborators ask most often.
Instruction-following review, factuality checking, hallucination detection, rubric-based scoring, and comparative response judgement. I work best on projects that value evidence-based, auditable ratings over speed alone.
Both. I build responsive, accessible sites from the ground up, and I can review and improve existing ones — from lightweight static pages to dashboards, support systems, and client portals. Scope is always right-sized to your needs.
Yes. I am based in Belfast, Northern Ireland, and I have spent years working remotely with distributed teams. Clear written communication and disciplined documentation keep everything on track.
Under Non-Disclosure Agreements where required. I document findings with evidence, impact, reproduction steps, and remediation guidance — without exposing sensitive details or unauthorised claims.
WhatsApp is fastest. Scan the QR code in the contact section or message +44 7882 753398. You can also email or connect on LinkedIn — I respond to all three.
Working Together
From first message to final delivery and beyond — a clear, predictable process designed to remove uncertainty.
We discuss your goals, timeline, and scope on WhatsApp, email, or LinkedIn. I ask clarifying questions and give honest feedback on feasibility.
Deliverable
Clear understanding of requirements
I send a written scope document with deliverables, timeline, and fixed price. No surprises — you know exactly what you're getting and what it costs.
Deliverable
Fixed quote + project brief
Work begins. For websites, you see progress early and can review direction. For AI evaluation, I apply the rubric methodology and document findings as I go.
Deliverable
Regular progress updates
Final delivery with documentation. For websites: deployment-ready code + handover notes. For evaluation: structured report with evidence and recommendations.
Deliverable
Final deliverable + docs
Post-delivery support is included. Questions, small tweaks, and clarifications are all part of the service — not an extra. Your success matters.
Deliverable
30 days included support
The first step is just a message. Tell me what you're working on and we'll figure out the rest together.
Get Started
Tell me about your website project. I'll review your requirements and respond on WhatsApp with next steps. No account needed.
Already started a discussion? Continue your website enquiry on WhatsApp.
Refer & Earn
Know someone who needs a professional website? Refer them and earn 10% when they become a qualifying paying client.
10%
Commission on qualifying projects
VC-7Q2Mwebsite-url/?ref=VC-7Q2MCodes are assigned manually after approval. The website does not validate whether a code is officially approved.
Available year-round. Reviewed manually.
Occasional updates on AI evaluation insights, prompt engineering patterns, and new project work. No spam — just thoughtful notes.
Contact
Available for AI evaluation, prompt engineering, web development, cybersecurity, technical consulting, collaboration, freelance projects, and employment opportunities.
Available for new work
AI evaluation, prompt engineering, web development, cybersecurity, technical consulting, collaboration, freelance projects, and employment opportunities.
Location
Belfast, Northern Ireland, United Kingdom