Systems I've Built
Browse active research, published Python packages, API clients and earlier engineering work. Each entry links to its source, documentation or live demo where available.
TrainLens
Evidence-grounded ML training diagnostics and next-experiment recommendations.
Turns notebook training state into interpretable diagnostics and performance visualizations. Deterministic local analysis produces actionable next steps; optional LLM reports and coding-agent integrations extend the workflow.
ASTScribe
Explain ML notebook methodology through static code and dependency analysis.
Analyzes supported Python ML code and Jupyter notebooks to identify model setup, training, inference and evaluation patterns. Traces cell dependencies and evidence to source lines, without running code or relying on an LLM.
LastLight
Offline retrieval from verifiable local sources, with evidence-based abstention.
A standard-library Python engine that searches Markdown and ZIP knowledge collections in low-connectivity environments, exposes source evidence and withholds answers without sufficient support. No cloud, embeddings or vector database required.
MiteCoder
CPU-first local coding assistant for resource-constrained workstations.
An early-stage offline coding agent that uses compact local models to inspect source files, apply workspace-scoped edits and validate them with configured checks. Offers a CLI and local web interface while GPUs are busy.
Semauri
Deterministic compilation from controlled natural language.
Explores predictable, human-readable programming constructs for generating web documents, JSON Schema, filesystem plans and typed ML workflows. Effects require explicit authorization; compilation itself has no side effects.
Cablegram
Token-efficient communication with auditable meaning preservation.
Research toolkit for minimizing token cost while retaining task-critical meaning using deterministic invariant checks, receiver-aware context, auditable candidate selection and benchmarks.
Suffice
Measure the smallest successful token budget for AI agents.
Explores token-efficiency frontiers across prompts, context, tools, memory and responses under explicit task-success constraints. Savings only count when outcomes remain successful.
MAVERICK
Inspectable multi-agent visual reasoning inspired by human cognition.
Four-agent perceive → describe → critique → refine loop for interpretable VLM analysis, uncertainty-aware reasoning and stronger evidence-grounded image descriptions.
Neural Audio Theory
Open technical guide to modern AI music generation.
Educational project covering signal processing, embeddings, transformers, diffusion architectures, training and prompt conditioning for developers and researchers.
music-to-text
Turn audio into structured music metadata and industry copy.
Local-first Python framework for acoustic feature extraction, reproducible audio analysis, structured exports, A&R notes, PR pitches, playlist descriptions and sync licensing copy. Works with local or OpenAI-compatible LLMs.
context-dedup
Deduplicate repeated LLM and agent context.
Reduces redundant contextual content in LLM- and agent-based workflows.
evidenceflow
Track propagation of evidence through AI-agent traces.
Deterministic analysis of how retrieved and tool-generated evidence is used across agent steps.
rag-chunk-audit
Audit chunk quality and segmentation in RAG pipelines.
Diagnostics for chunk segmentation and retrieval quality during RAG development.
embedding-drift-lite
Detect shifts in embedding distributions.
Lightweight embedding drift detection and inspection for retrieval and ML evaluation.
parametricbench
Provider-independent LLM/VLM regression benchmarking.
Compare model and system behavior across providers to identify regressions.
promptshield-llm
Screen risky or injected input in LLM pipelines.
Prompt-injection, unsafe-instruction and input-risk screening for language-model applications.
metaclean-vlm
Normalize metadata for image and VLM datasets.
Dataset metadata cleaning and normalization for multimodal workflows.
visual-patch-audit
Inspect image patches and visual signals for VLM tasks.
Patch-level auditing tools for diagnosing multimodal visual inputs.
vlm-occlusion
Probe visual-claim sensitivity with grid occlusion.
Black-box experiments that occlude image regions to assess how VLM claims change.
vlm-prior-probe
Test whether VLMs follow images or learned priors.
Counterfactual black-box evaluation of visual evidence versus prior-driven responses.
matplotlib-dark
Dark themes for Matplotlib with safe temporary styling.
Automatic dark-mode plotting, ready-to-use themes and temporary styles.
egypttranslit
Egyptological transliteration to clean Unicode.
Python/CLI conversion from MdC notation with normalization, validation and diagnostics.
text-to-music-prompt-structurer
Extract structured prompts for text-to-music generation.
Structured prompt extraction and organization for generative-audio workflows.
llm7R
LLM7.io client for R analysis and multimodal workflows.
Lightweight R client supporting chat, streaming, model discovery, JSON mode, tool calling, data-frame analysis, vision and image/video generation.
llm-ts-api-wrapper
Zero-runtime-dependency TypeScript client for OpenAI-compatible APIs.
Supports chat and Responses APIs, streaming, embeddings, model discovery, tool calls, retries, timeouts and typed errors. Designed for source integration; not published on npm.
edujbarrios-ui
Reusable frontend components for AI tools and interfaces.
Component library for clean AI demos, evaluation views, model outputs, agent states and developer-oriented products.
AirLLM for VLM
Memory-efficient VLM inference on constrained hardware.
Explores layer streaming and model weight offloading to run multimodal models within limited GPU memory.
AutoPromTune
Automatically refine underspecified prompts with LLMs.
Multi-pass prompt editing and alternative generation to turn vague requests into more usable instructions.
MedGemma Agentic Workflow
Experiments in reproducible agentic medical-AI workflows.
Research framework for defined clinical AI tasks and agent boundaries using foundation models.
Music Creator Agent
Cooperating agents for structured AI music prompts.
Genre, tone, lyric and prompt-specialist agents explore more deliberate music-generation workflows.
C-Notebook (cnb)
Cell-based interactive execution for C.
Experimental notebook engine for sequential C compilation, persistent state and plain-text notebooks.
ncmds
Zero-configuration Markdown documentation site generator.
Markdown-first CLI that turns existing documentation into a navigable site with minimal setup.
FEMOG
Domain-indexed directory of engineering GitHub profiles.
Open curation project for discovering engineering work through structured domains and contributions.
NEONMIX
Browser-native stem mixing and mastering.
Client-side Web Audio API experiments for multistem audio processing without a DAW.
SafeID
Privacy-first, zero-upload document watermarking.
Sensitive documents are processed entirely in the browser, without server uploads.
Voice-To-Markdown
Voice-controlled Markdown authoring.
Bilingual, structured Markdown drafting using browser speech recognition.
llm7.io GUI
Browser-based LLM inference and prompt testing.
Lightweight browser UI to test completions and prompts without building a custom client.