Crit by colorME
×
Design Thinking, Agile MVP and a pilot plan for an AI Portfolio Coaching Assistant at colorME, with a working product.



Portfolios published by Vietnamese designers, 2025–26
Feedback arrives after they stopped working on the piece. Weaker portfolio at course end.Assumption
20–40 pieces × 10–15 min ≈ 3–10 h of evening critique per assignment round.Estimate
Portfolio quality drives completion, referrals and upsell into longer tracks (up to 20.3M VND).Fact · colorme.vn






And this is what the year has to produce: 2025–26 portfolios published on Behance by Vietnamese designers. Portfolio quality is the outcome; it is built one critique cycle at a time.

Marketing exec, Photoshop class in the evening, wants an in-house design job within a year.
Fear of AI: generic feedback; being judged by a machine.

Senior agency designer teaching two evening classes.
Fear of AI: bland, over-confident feedback that undermines him.
Feedback must land while the file is still open.
The skill we teach is seeing your own work. Train the eye, don't fix the piece.
Where, what, why, tied to a rubric criterion.
Trust needs visible limits: no generated designs, instructor override.



No AI feedback until they state their intent and rate themselves on every rubric criterion. The gap between their rating and the crit is where learning happens.
Pinned observation, one Socratic question and one principle per criterion. No generated images, rewritten copy or exact specs. “Do it for me” gets redirected.
Final work, low-confidence crits and self-vs-AI gaps ≥ 2 levels go to the instructor, who can agree, adjust or replace every AI judgement.
Synthetic learner and poster. Real model output.

Learner: v1 → v2 with rubric deltas (“Developing → Proficient”). Average +1.2 levels on this synthetic example.

Instructor: confidence per criterion; agree / adjust / replace; note; timed review. Vietnamese crit shown.
Reads the artwork (IDs 26, 22, 23), writes per-criterion feedback as JSON (27, 34), personalises to intent and language (20), answers and classifies chat (30, 32).
Excluded: ID 33, image generation.
Reflect-first gate · confidence floor 0.6 · self-vs-AI gap ≥ 2 → hold · final → hold · prescriptive-output filter · daily limits · role-based access · audit trail.
Instructor: rubric, held crits, overrides, final grades. Learner: every design decision. Academic lead: go-live, prompt changes, pause.
| # | Assumption | Cheapest test | Pass if |
|---|---|---|---|
| 1 | AI ≈ instructor judgement | Golden set: 60 past pieces, 2 instructors, blind | ±1 level on ≥ 85% |
| 2 | Learners reflect first | Funnel: upload → reflection | ≥ 80% complete |
| 3 | Faster feedback → more, better revisions | Treatment vs control classes | ≥ 60% resubmit; +0.5 level |
| 4 | Instructors save time | Timed reviews + retro | −40% min/learner |
| 5 | Coach never does the work | Red-team prompts; audit 50 crits | 0 prescriptive; 100% deflected |
| 6 | Learners consent | Consent screen, week 1 | ≥ 90% opt in |
One assignment (event poster), one 5-criterion rubric · reflect-first · pinned crit EN/VI · routing + instructor review · plan → v2 compare · guarded chat · pilot dashboard
AI grades (trust, fairness) · generated designs (does the work) · LMS integration (40 learners enrol by link) · other media (prove one first) · cohort digest (needs data)
Cupcake · pilot6 weeks · 1 assignment · 2 classes
Birthday cake · foundations2 quarters · all assignments
Wedding cake · whole school12 months +Delivery: 2-week sprints · academic lead = product owner · 1 developer · both pilot instructors in every sprint review · evidence gates, not dates.
All five hit · no incident · instructors want to continue
1–2 miss, learning ≥ neutral → fix, one more cohort
Disagreement > 35% · learning below control · any privacy incident
| Risk | Control |
|---|---|
| Confident, wrong feedback | Confidence floor + gap routing · visible confidence · “I disagree” · weekly calibration |
| Over-reliance | Reflect-first · questions not answers · plan in own words · no generated designs |
| Everyone designs “for the AI” | Level 4 rewards originality · instructors grade finals |
| Instructors feel replaced | They author the rubric, override everything, grade finals |
| Privacy, minors | Decree 13/2023 consent · guardian consent · paid tier, no training · no names in prompts |
| Silent model drift | Pinned model · golden-set re-run on any change · fallback · pause switch |
Govern named owners, change control · Map beginners, minors, IP · Measure scorecard by language & style · Manage routing, override, pause
Instructors co-write rubrics, see golden-set results first, 1-h onboarding, a champion. Learners “the AI can be wrong, disagree” from session 1. Academic lead owns a weekly 30-min calibration ritual.
A paid AI key with no training on data · a named academic owner · 2 instructors × ~6 h calibration · 2–3 dev weeks to add consent and a golden-set runner
≈ US$0.005–0.02 per critEstimate · under ~5,000 VND per learner per course · hosting in Cloudflare's free/entry tier
Golden-set agreement < 75% · learners skip reflection · learning below control · instructors spend more time
Fast feedback is easy now. Feedback that builds the learner's own eye is the product.
Try it: colorme-crit.pages.dev · Code: github.com/hungkhi/colorme-crit · Appendix follows
×
| ID | Capability | Use in Crit |
|---|---|---|
| 26 | Image captions | Describe elements & relationships |
| 22 | Identify objects | Locate elements → pins |
| 23 | Image → text | Read copy incl. Vietnamese diacritics |
| 27 | Image metadata | Level, confidence, box as JSON |
| 34 | Generate text | Observation, question, next step |
| 20 | Personalize | Intent, language, calibration, plan check |
| 30 · 32 | Q&A · intent | Coach chat; do-my-work detection |
| 19 | Moderate | Prescriptive-output filter |
| 04 · 28 | Patterns · summarize | Next: cohort digest |
| 33 | Generate images | Excluded by design |
Non-AI fixes (self-check, peer circles, videos) help and are included, but none gives specific feedback on this piece within minutes. Formative feedback tolerates labelled, contestable error; grades don't, so grades stay human.
Formative crit = Recommend · Final work = Draft, human approves · Never acts alone. Climb only when errors are reversible, measurable and owned.
| Item | Specification |
|---|---|
| Inputs | Artwork (≤ 8 MB, downscaled to 2,000 px) · brief · rubric (criteria × 4 descriptors) · intent (≥ 20 chars) · focus · self-rating 1–4 on every criterion · EN/VI · final flag · v2+: previous levels + plan |
| Outputs | Per criterion: level 1–4, confidence 0–1, observation, question, next step, box [ymin,xmin,ymax,xmax] on 0–1000 · 2 strengths · 1 priority · 1 reflection prompt · JSON-schema constrained |
| Routing | Hold if final · any confidence < 0.6 · |self − AI| ≥ 2 · missing criterion. Else release. Thresholds are pilot parameters. |
| Guardrails | No image generation in the product · prompt forbids fonts/hex/sizes/copy/layouts · post-filter rewrites prescriptive steps · chat intent classifier |
| Model | Gemini Flash vision + fallback model, 2 attempts each; prompt and schema are model-agnostic |
| Constraints | ≤ 20 crits & 60 chat messages / learner / day · p90 < 60 s · < US$0.02 / crit |
| Acceptance | Golden set ≥ 60 double-marked pieces: ±1 level ≥ 85%, exact ≥ 60% · pins correct ≥ 80% · 0 prescriptive in 50-crit audit · 100% do-my-work deflected · EN/VI gap ≤ 5 pts |
| Layer | Metric | Threshold |
|---|---|---|
| Business | Rubric improvement v1 → final vs control · resubmission · completion | ≥ +0.5 · ≥ 60% · no drop |
| UX | Reflection completion · comments rated helpful · would use again | ≥ 80% · ≥ 70% · ≥ 75% |
| Output quality | Agreement with instructor consensus (±1 / exact) · pin accuracy | ≥ 85% / 60% · ≥ 80% |
| Safety | Prescriptive outputs in audit · do-my-work deflection | 0 · 100% |
| Fairness | Agreement gap EN vs VI; hand-drawn vs digital; Vietnamese vs Latin type | ≤ 5 pts |
| Robustness | Low-res, blank, non-design, text-heavy uploads | Low confidence → held |
| Latency · cost | p90 crit time · cost per crit (stored per crit) | < 60 s · < US$0.02 |
Upload & reflect → authorize → critique → route → instructor review → release & plan
Rubric / prompt / model change → version → re-run golden set → academic lead approves → release
Build workflow & experience · Buy the model as an API · Reuse Google identity, Cloudflare · Crosses the boundary: artwork + intent to Gemini → paid tier + consent in production.
| Pri. | Story | Status |
|---|---|---|
| Must | Rate myself before feedback | Done |
| Must | Comments pinned & tied to criteria | Done |
| Must | Uncertain/final crits come to instructor; override any level | Done |
| Must | Plan → v2 → see what moved | Done |
| Must | Consent (incl. guardians) before first upload | Sprint 1 |
| Must | Golden-set runner gates every change | Sprint 1 |
| Should | Rubric versioning · queue filters | Sprint 2 |
| Should | Weekly cohort digest | Sprint 3 |
| Won't | AI grades · AI-generated designs | By design |
| Dimension | With Crit |
|---|---|
| Experience | Private, pinned crit in minutes; questions, not verdicts; visible progress |
| Process | Reflect → crit → plan → v2 → instructor-graded final |
| People | Instructors author rubrics & review exceptions; academic lead owns quality |
| Data | Every version, rating, override and reaction structured and measurable |
| Operations | Weekly calibration, dashboards, pause switch, manual fallback |

Live dashboard computed from the database. Synthetic demo data shown, n = 3 submissions.
| Item | Pilot | Scale / yr |
|---|---|---|
| AI inference | ~400 crits, under US$10 | $0.005–0.02 per crit |
| Hosting | Free tier | ~US$5–50 / mo |
| Development | 2–3 person-weeks | 0.5 FTE |
| Instructor time | ~12 h calibration | ~2 h / week |
Optimise cost per successful outcome (a learner +1 level on a criterion), not cost per API call. Token prices to be re-verified before scale.
Fact verified (colorme.vn, prototype runs) · Estimate calculated · Assumption founder's view · Hypothesis pilot tests it. No interviews or results are presented as real. All demo accounts, posters and metrics are synthetic.
No primary research yet (week 0 adds 5 learner + 2 instructor interviews) · model confidence is self-reported, so the 0.6 floor must be calibrated · prototype uses a free-tier AI key with synthetic data only.
2025–26 portfolios published on Behance by Vietnamese designers. A colorME learner is measured against this when they apply for work — which is why feedback cycles per piece, not hours taught, is the number that matters.



















Covers reproduced for academic illustration; all rights remain with their authors (full credits in the report, §18). None of these designers is claimed as a colorME learner.