Open roles across every project and research thread.
Harden and extend the offline assistant — improve how reliably small on-device models handle assessment, tutoring, and content generation, and validate the outputs against curriculum standards and real classroom needs. Good if you like making agent systems actually hold up on constrained hardware.
Build the part of the system that enforces hard safety rules and keeps its decisions traceable. Good if you like LLM agents and building things that fail safe.
Build and improve the full grading pipeline end to end — agent logic, document ingestion, RAG retrieval, and rubric-aligned generation. Measure quality with benchmarking and ablation studies, with a path to co-authoring on a live institutional AI deployment. Good if you want to go deep on a system that runs in real university courses.
Extend the multi-agent retrieval pipeline — improve specialist routing, retrieval quality, and LLM-scored ranking. Work directly on a live system used for clinical evidence search. Good if you want to go deep on agentic RAG beyond toy examples.
Extend and improve the multi-agent pipeline — refine taste vector matching, add specialist agents (dietary, cuisine-specific), and improve group preference aggregation. Good if you want to go deep on multi-agent systems and personalization with real users.
Evaluate retrieval quality across clinical domains, build benchmark datasets, and run ablation studies on retrieval and ranking strategies. The project has an active research thread — path to co-authorship. Good if you like measuring whether a system actually works and care about evidence quality in healthcare.
Applied Scientist — Curriculum & Evaluation
SeeSay →Turn the Graded Direct Method into a real graded curriculum and measure whether learners actually acquire the language.
Bring the linguistic judgment: old scripts, dead and low-resource languages, and what a correct translation actually means here.
Train the models that decide what gets recommended — retrieval, ranking, and sequential modeling. Every change is measured against a real offline eval harness (Recall@K, NDCG) before it ships. Good if you want to build recsys the way large-scale systems are built — not just cosine similarity.
Design the study that shows this works, and extend it to a second domain.
CCI framework architecture, methodology coordination, and paper positioning.
Build the data and delivery layer behind the product: a Postgres schema for users, flavor profiles, menus, and a persisted safety-gate audit trail, with the full test suite and eval benchmarks gating every merge.
Build and maintain the product end to end: the FastAPI backend, SSE streaming layer, and AWS infrastructure, plus the clinician-facing search interface — real-time SSE evidence streaming, source highlighting, and quality indicators. Keep a live clinical tool fast, reliable, and observable.
Build the product end to end — FastAPI backend, WebSocket group chat, social graph (friends, groups, collections), AWS Cognito auth, and the React discovery feed UI. This is a social product so both sides matter: fast APIs and a UX that makes group dining decisions fast and clear.
Help build and maintain the AnacodicAI Labs website — the site you're looking at right now.
Owns the first-author writing and synthesis — turn the section drafts into one coherent argument, own the figures and the framing. Good if you want a lead-author byline on an invited review.
Own the system end to end — matching, coordination, and making it work on poor connectivity.
Leads benchmarking design, measurement framework, and experimental execution.
Generate and verify the picture panels — one new element per step, and the image has to actually mean what the sentence says.
Own calibration and analysis — measure whether a learner's judgment is genuinely improving, not just their confidence.
Own recommendations and longitudinal evaluation — is this genuinely helping, or just adding noise?
Own the OCR stage and its evaluation — catching the errors that a fluent translation would otherwise hide.
Machine Learning Engineer — Speech
SeeSay →Build the voice and pronunciation loop: clear speech out, learner speech in, useful feedback back.
Build the benchmark that shows how close the offline system gets to expensive cloud models — quality across the three tasks, plus the cost, latency, and memory of running on low-cost edge hardware. Good if you like building the arena that makes a result credible.
Build the multi-region benchmark and the baseline suite the method must beat. Good if you like building the arena that makes a result credible.
Build the tests that prove the system is safe, not just accurate — design the evaluation and run the comparisons.
Build the part that checks how consistent a questioned work is with an artist's confirmed self-portraits, and the calibrated confidence layer that can say 'not enough evidence' instead of guessing. Good if you like metric learning and honest probabilities.
Build the vision pipeline that turns paintings into color and visual features, and the analysis and statistical test that check whether the changes across a career line up with real events better than chance. Good if you like turning pixels into defensible results.
Open a second deployment vertical: new producers, new buyers, same honest measurement.
Help gather the real-world data that shows where SafeBite holds up and where it falls short — running small pilots with the right campus programs (we'll connect you) and bringing structured feedback back to the team. Good if you like turning messy real-world use into signal people can build on.
Account for the energy, cost, and carbon of the system per query, and quantify its cost-versus-benefit tradeoff. Good if you care about the carbon of the intelligence, not just its accuracy.
Assemble the Van Gogh and Munch painting corpus from museum open data and curate the timeline of life events the study draws on. Good if you like building clean, well-sourced datasets.
Build the provenance and sources retrieval and the layer that combines the different kinds of evidence into one cited rationale. Good if you like grounded, auditable systems.
Owns the research question and the paper end to end, with faculty mentorship.
Help survey and synthesize part of the literature into the shared draft — how the footprint is measured, where efficiency and scheduling gains come from, or what honest accounting requires. Good if you like turning a pile of papers into a clear argument.
Build the product end to end: the practice loop, the session UI, and the APIs behind it.
Build the time-budget and constraint layer — the part that guarantees study time is never quietly eaten.
Build the platform: APIs, matching, and an interface that works on a low-end phone.
Put SafeBite through its paces with real menus and real dietary needs, log what breaks or falls short, and turn it into structured data the team uses to make it safer. Good if you like finding gaps and writing them up clearly.