Crit by colorME
Nguyen Viet Hung · VinUni MBA
colorME×VinUniversity
Individual Project · Scenario 18

An AI coach that makes design learners look again.

Design Thinking, Agile MVP and a pilot plan for an AI Portfolio Coaching Assistant at colorME, with a working product.

PresenterNguyen Viet Hung · Founder, colorME
CourseAI Product Design & Innovation · MB21CI11
Live prototypecolorme-crit.pages.dev
InstructorDr. Le Nhan Tam
Portfolios published by Vietnamese designers, 2025–26
01 The problem

Learners get one feedback cycle, where design skill needs three or four.

Briefin classMakealone, lateSubmitone exportWait2–4 daysCrita few piecesMove onrevision is rare
Learners

Can't say what's wrong

Feedback arrives after they stopped working on the piece. Weaker portfolio at course end.Assumption

Instructors

Same five notes, thirty times

20–40 pieces × 10–15 min ≈ 3–10 h of evening critique per assignment round.Estimate

colorME

Feedback is capped by instructor hours

Portfolio quality drives completion, referrals and upsell into longer tracks (up to 20.3M VND).Fact · colorme.vn

And this is what the year has to produce: 2025–26 portfolios published on Behance by Vietnamese designers. Portfolio quality is the outcome; it is built one critique cycle at a time.

02 Empathize proto-personas · validate in week 0

Two people, two fears.

Mai Anh, 24 · the career switcher

Marketing exec, Photoshop class in the evening, wants an in-house design job within a year.

Says“Is this okay?” · “I'll fix it later.”
ThinksMy colours are bright, so it must stand out.
DoesTries 4 fonts, posts once, rarely revises.
FeelsUnsure; exposed in public crits.

Fear of AI: generic feedback; being judged by a machine.

Anh Tuấn, 31 · the instructor

Senior agency designer teaching two evening classes.

Says“Squint at it: what do you read first?”
ThinksMost issues are the same five.
DoesCritiques at night; deep-dives a few pieces.
FeelsTired of repetition; protective of his voice.

Fear of AI: bland, over-confident feedback that undermines him.

02 Reflect · as-is journey

The pain is in the gaps: stuck, forgotten, exposed.

Brief
Make
Submit
Wait
Feedback
Move on
Doing
Listens in class
Works alone; tries many fonts
Posts one export
Starts next exercise
Hears 1–2 generic comments
Keeps v1 for portfolio
Thinking
“Seems easy.”
“Something's off, but what?”
“Hope it's okay.”
“Did he see it?”
“What exactly should I change?”
“No time to redo it.”
Feeling
Motivated
Stuck
Anxious
Forgotten
Exposed
Resigned

Speed

Feedback must land while the file is still open.

Self-critique

The skill we teach is seeing your own work. Train the eye, don't fix the piece.

Anchored

Where, what, why, tied to a rubric criterion.

Boundaries

Trust needs visible limits: no generated designs, instructor override.

Make — alone, after class
Crit — a few pieces in depth, in public
Move on — before anything is revised
02 Define · Hills
HILL 1 · LEARNERA beginner design learner NEEDS TO know, the same evening, what to change and why, BECAUSE feedback days later arrives after they stopped working on it.
HILL 2 · INSTRUCTORAn instructor NEEDS TO spend critique time where their judgement changes the outcome, BECAUSE the same first-pass note thirty times leaves no time for learners who need them.
HILL 3 · ACADEMYcolorME NEEDS TO give every learner several feedback cycles without adding instructor hours, BECAUSE portfolio quality drives completion and referrals.
Problem statement

How might we give every learner specific, rubric-grounded feedback within minutes, in a way that strengthens their own eye, while instructors keep control?

03 Ideate & prioritize

Twelve ideas. The two easiest ones were unwise.

1Peer-crit circles with rubric
2Rubric self-check
3Common-mistakes video library
4AI auto-grader, score only
5AI coach: pinned notes + questions, reflect-first
6AI generates the “improved” poster
7Instructor voice notes
8Weekly live group crit
9Alumni mentor marketplace
10Cohort mistake digest
11Version compare v1 → v2
12Guarded coach chat
BIG BETSNO-BRAINERSSAVE FOR LATER / UNWISEUTILITIESHIGH IMPACTLOWHARDEASY → FEASIBILITY9 Mentor marketplace10 Cohort mistake digest5 AI coach, reflect-first11 Version compare12 Guarded chat2 Self-check6 AI redesign4 Auto-grader1 Peer circles3 Video library7 Voice notes8 Live crit
03 The concept

Crit: three rules that make fast feedback safe for beginners.

1

The learner judges first

No AI feedback until they state their intent and rate themselves on every rubric criterion. The gap between their rating and the crit is where learning happens.

2

The AI coaches, never designs

Pinned observation, one Socratic question and one principle per criterion. No generated images, rewritten copy or exact specs. “Do it for me” gets redirected.

3

The instructor decides what matters

Final work, low-confidence crits and self-vs-AI gaps ≥ 2 levels go to the instructor, who can agree, adjust or replace every AI judgement.

Reflectintent + self-ratingCrit in ~1 minpinned, rubric-tiedRouterelease or holdPlan“what I’ll change”v2see what moved
04 Prototype · live at colorme-crit.pages.dev
Crit on v1

Specific, pinned, and honest about uncertainty.

  • Named the real issues: date competing with the title, three typefaces, low-value tagline, clip-art concept.
  • Each comment is pinned to the element.
  • AI level shown against the learner's own: colour 4 vs 2 → held for instructor.
  • ≈ 10 s per crit, English or Vietnamese.Fact · test run

Synthetic learner and poster. Real model output.

04 Prototype · the loop

Plan → v2 → see what moved. The instructor sees what matters.

v1 v2 compare

Learner: v1 → v2 with rubric deltas (“Developing → Proficient”). Average +1.2 levels on this synthetic example.

Instructor review

Instructor: confidence per criterion; agree / adjust / replace; note; timed review. Vietnamese crit shown.

04 How it works

A workflow, not an agent. AI assists · systems enforce · people decide.

AI assists

Reads the artwork (IDs 26, 22, 23), writes per-criterion feedback as JSON (27, 34), personalises to intent and language (20), answers and classifies chat (30, 32).

Excluded: ID 33, image generation.

Systems enforce

Reflect-first gate · confidence floor 0.6 · self-vs-AI gap ≥ 2 → hold · final → hold · prescriptive-output filter · daily limits · role-based access · audit trail.

People decide

Instructor: rubric, held crits, overrides, final grades. Learner: every design decision. Academic lead: go-live, prompt changes, pause.

01–02Experience & appResponsive web app: learner, instructor, pilot dashboard · Google sign-in
03–04Orchestration & AIclaim → critique (Gemini Flash vision, JSON schema, fallback model) → policy routing → release
05–07Data & infraInstructor-owned rubrics · artwork in R2 · crits, reviews, reactions in D1 · Cloudflare Pages Functions · CI/CD from GitHub
05 Agile · assumptions & MVP

Test the riskiest assumption first: does it agree with instructors?

#AssumptionCheapest testPass if
1AI ≈ instructor judgementGolden set: 60 past pieces, 2 instructors, blind±1 level on ≥ 85%
2Learners reflect firstFunnel: upload → reflection≥ 80% complete
3Faster feedback → more, better revisionsTreatment vs control classes≥ 60% resubmit; +0.5 level
4Instructors save timeTimed reviews + retro−40% min/learner
5Coach never does the workRed-team prompts; audit 50 crits0 prescriptive; 100% deflected
6Learners consentConsent screen, week 1≥ 90% opt in

MVP: in (built)

One assignment (event poster), one 5-criterion rubric · reflect-first · pinned crit EN/VI · routing + instructor review · plan → v2 compare · guarded chat · pilot dashboard

Out, and why

AI grades (trust, fairness) · generated designs (does the work) · LMS integration (40 learners enrol by link) · other media (prove one first) · cohort digest (needs data)

05 Agile · delivery & learning loop

Every week: learn from disagreement, re-test, then release.

Signalsdisagree · overridesCalibrate20 crits weeklyAdjustrubric · promptRe-testgolden + red-teamReleaselead approves
Cupcake · pilot6 weeks · 1 assignment · 2 classes
Birthday cake · foundations2 quarters · all assignments
Wedding cake · whole school12 months +
Experience
“A real crit on my poster the same evening.”
“Crits on every assignment; my instructor sees where the class is stuck.”
“From first poster to job-ready portfolio, I can explain my growth.”
Adds
Reflect-first crit, routing, review, versions, chat
Logo & social rubrics · cohort digest · rubric versioning · LMS sync · paid AI tier
UI/UX & motion crits · portfolio review before job applications · B2B training

Delivery: 2-week sprints · academic lead = product owner · 1 developer · both pilot instructors in every sprint review · evidence gates, not dates.

06 Pilot

2 classes vs 2 control classes · one poster · six weeks.

< 10 min
median feedback turnaround
baseline 2–4 days
≥ 60%
resubmit after feedback
vs control
+0.5
rubric level v1 → final, over control
blind-graded finals
−40%
instructor minutes per learner
timed reviews
≤ 20%
AI–instructor disagreement
overrides / criteria
W0W1W2W3W4W5Golden set, red-team set, consent flowInstructor onboarding (1 h) + rubric co-writingLive pilot: crits, reviews, versionsWeekly calibration + sprint reviewFinal grading (blind), analysisScale / adjust / stop decisionSet-up & decision gatesLive pilotRecurring / supporting

Scale

All five hit · no incident · instructors want to continue

Adjust

1–2 miss, learning ≥ neutral → fix, one more cohort

Stop

Disagreement > 35% · learning below control · any privacy incident

06 Risks, Responsible AI & change

The real risks are human, so the controls are too.

RiskControl
Confident, wrong feedbackConfidence floor + gap routing · visible confidence · “I disagree” · weekly calibration
Over-relianceReflect-first · questions not answers · plan in own words · no generated designs
Everyone designs “for the AI”Level 4 rewards originality · instructors grade finals
Instructors feel replacedThey author the rubric, override everything, grade finals
Privacy, minorsDecree 13/2023 consent · guardian consent · paid tier, no training · no names in prompts
Silent model driftPinned model · golden-set re-run on any change · fallback · pause switch
NIST AI RMF

Govern named owners, change control · Map beginners, minors, IP · Measure scorecard by language & style · Manage routing, override, pause

Change management

Instructors co-write rubrics, see golden-set results first, 1-h onboarding, a champion. Learners “the AI can be wrong, disagree” from session 1. Academic lead owns a weekly 30-min calibration ritual.

07 Decision

Approve a 6-week controlled pilot: 2 instructors, ~40 learners, one assignment.

We need

A paid AI key with no training on data · a named academic owner · 2 instructors × ~6 h calibration · 2–3 dev weeks to add consent and a golden-set runner

It costs

≈ US$0.005–0.02 per critEstimate · under ~5,000 VND per learner per course · hosting in Cloudflare's free/entry tier

We'd stop if

Golden-set agreement < 75% · learners skip reflection · learning below control · instructors spend more time

Fast feedback is easy now. Feedback that builds the learner's own eye is the product.

Try it: colorme-crit.pages.dev · Code: github.com/hungkhi/colorme-crit · Appendix follows

colorME×VinUniversity
Appendix A1 capabilities

AI suitability & capability map

IDCapabilityUse in Crit
26Image captionsDescribe elements & relationships
22Identify objectsLocate elements → pins
23Image → textRead copy incl. Vietnamese diacritics
27Image metadataLevel, confidence, box as JSON
34Generate textObservation, question, next step
20PersonalizeIntent, language, calibration, plan check
30 · 32Q&A · intentCoach chat; do-my-work detection
19ModeratePrescriptive-output filter
04 · 28Patterns · summarizeNext: cohort digest
33Generate imagesExcluded by design

Is it really an AI problem?

Non-AI fixes (self-check, peer circles, videos) help and are included, but none gives specific feedback on this piece within minutes. Formative feedback tolerates labelled, contestable error; grades don't, so grades stay human.

Autonomy ladder

Formative crit = Recommend · Final work = Draft, human approves · Never acts alone. Climb only when errors are reversible, measurable and owned.

Appendix A2 spec

AI product specification

ItemSpecification
InputsArtwork (≤ 8 MB, downscaled to 2,000 px) · brief · rubric (criteria × 4 descriptors) · intent (≥ 20 chars) · focus · self-rating 1–4 on every criterion · EN/VI · final flag · v2+: previous levels + plan
OutputsPer criterion: level 1–4, confidence 0–1, observation, question, next step, box [ymin,xmin,ymax,xmax] on 0–1000 · 2 strengths · 1 priority · 1 reflection prompt · JSON-schema constrained
RoutingHold if final · any confidence < 0.6 · |self − AI| ≥ 2 · missing criterion. Else release. Thresholds are pilot parameters.
GuardrailsNo image generation in the product · prompt forbids fonts/hex/sizes/copy/layouts · post-filter rewrites prescriptive steps · chat intent classifier
ModelGemini Flash vision + fallback model, 2 attempts each; prompt and schema are model-agnostic
Constraints≤ 20 crits & 60 chat messages / learner / day · p90 < 60 s · < US$0.02 / crit
AcceptanceGolden set ≥ 60 double-marked pieces: ±1 level ≥ 85%, exact ≥ 60% · pins correct ≥ 80% · 0 prescriptive in 50-crit audit · 100% do-my-work deflected · EN/VI gap ≤ 5 pts
Appendix A3 evaluation

Multi-layer evaluation scorecard

LayerMetricThreshold
BusinessRubric improvement v1 → final vs control · resubmission · completion≥ +0.5 · ≥ 60% · no drop
UXReflection completion · comments rated helpful · would use again≥ 80% · ≥ 70% · ≥ 75%
Output qualityAgreement with instructor consensus (±1 / exact) · pin accuracy≥ 85% / 60% · ≥ 80%
SafetyPrescriptive outputs in audit · do-my-work deflection0 · 100%
FairnessAgreement gap EN vs VI; hand-drawn vs digital; Vietnamese vs Latin type≤ 5 pts
RobustnessLow-res, blank, non-design, text-heavy uploadsLow confidence → held
Latency · costp90 crit time · cost per crit (stored per crit)< 60 s · < US$0.02
Appendix A4 architecture

Eight layers · two flows · build/buy/reuse

01ChannelsResponsive web app; later embedded in colorme.vn homework pages
02Business appsAssignments, submissions, versions, reviews, responses; later LMS sync
03OrchestrationDeterministic workflow: claim → critique → route → release
04AI servicesGemini Flash vision + JSON schema; fallback; policy module
05Data & knowledgeInstructor-owned rubrics & briefs; R2 artwork; D1 records
06IntegrationGoogle OAuth (PKCE) · Gemini API · GitHub Actions → Cloudflare
07InfrastructureCloudflare Pages + Functions, D1, R2
08Security & governanceHashed HttpOnly sessions · roles · per-learner file isolation · rate limits · audit trail · pause switch

Serve users

Upload & reflect → authorize → critique → route → instructor review → release & plan

Keep current

Rubric / prompt / model change → version → re-run golden set → academic lead approves → release

Build workflow & experience · Buy the model as an API · Reuse Google identity, Cloudflare · Crosses the boundary: artwork + intent to Gemini → paid tier + consent in production.

Appendix A5 backlog & change

Prioritized backlog · what changes

Pri.StoryStatus
MustRate myself before feedbackDone
MustComments pinned & tied to criteriaDone
MustUncertain/final crits come to instructor; override any levelDone
MustPlan → v2 → see what movedDone
MustConsent (incl. guardians) before first uploadSprint 1
MustGolden-set runner gates every changeSprint 1
ShouldRubric versioning · queue filtersSprint 2
ShouldWeekly cohort digestSprint 3
Won'tAI grades · AI-generated designsBy design
DimensionWith Crit
ExperiencePrivate, pinned crit in minutes; questions, not verdicts; visible progress
ProcessReflect → crit → plan → v2 → instructor-graded final
PeopleInstructors author rubrics & review exceptions; academic lead owns quality
DataEvery version, rating, override and reaction structured and measurable
OperationsWeekly calibration, dashboards, pause switch, manual fallback
Appendix A6 dashboard & cost
Pilot dashboard

Live dashboard computed from the database. Synthetic demo data shown, n = 3 submissions.

Cost range Estimate

ItemPilotScale / yr
AI inference~400 crits, under US$10$0.005–0.02 per crit
HostingFree tier~US$5–50 / mo
Development2–3 person-weeks0.5 FTE
Instructor time~12 h calibration~2 h / week

Optimise cost per successful outcome (a learner +1 level on a criterion), not cost per API call. Token prices to be re-verified before scale.

Appendix A7 evidence & AI use

Evidence, limitations and AI-use disclosure

Evidence tags

Fact verified (colorme.vn, prototype runs) · Estimate calculated · Assumption founder's view · Hypothesis pilot tests it. No interviews or results are presented as real. All demo accounts, posters and metrics are synthetic.

Limitations

No primary research yet (week 0 adds 5 learner + 2 instructor interviews) · model confidence is self-reported, so the 0.6 floor must be calibrated · prototype uses a free-tier AI key with synthetic data only.

AI-use disclosure (Level 4, syllabus §7)

  • Claude Code (Claude Opus 5): read the course materials on Canvas to align structure and vocabulary; drafted the report and this deck; built, tested and deployed the prototype on my direction.
  • Gemini Flash: the AI inside the product (critique, coach chat).
  • HTML + headless Chrome: synthetic demo posters.
  • Checks: type checks, 7 unit tests (routing, overrides, guardrail), live end-to-end runs, CI on every push; every colorME number tagged by evidence type.
  • I remain accountable for every claim and decision. Details in report §17.
Appendix A8 why portfolios

The artefact every course builds towards

2025–26 portfolios published on Behance by Vietnamese designers. A colorME learner is measured against this when they apply for work — which is why feedback cycles per piece, not hours taught, is the number that matters.

Ngọc Khuê
Asae Phạm
Han Giaaa
Hoàng Phúc I HOPU
Gia Khiem Pham
Khanh Nimal
Trung Nghia
Linh Lâm
Nghĩa Ming
Ha Thanh
Typha Nguyen
Hải Nguyễn
Hạnh Minh Phạm
Minh Chau
Mike Summer
HUYNH NHU
Tuan Nguyen
Huyền Phương
LÝ LAM

Covers reproduced for academic illustration; all rights remain with their authors (full credits in the report, §18). None of these designers is claimed as a colorME learner.