ML-Ticker
FactSet FQL CLI with a restricted formula language and DAG-ordered metrics. Merrill Lynch internship; private repo.
Built
Python CLI for equity research: FactSet FQL snapshots or timeseries, YAML/Pydantic configs, a sandboxed formula engine, and Rich/Excel export. A planner batches FQL calls, then Python metrics run in topological order. FDT mode scores a name against peers into Bull/Base/Bear multiples.
Hard parts
- Formulas needed to be useful without becoming a code-execution hole (AST denylist + allowlisted builtins).
- Interdependent metrics: DAG, cycle path, typo suggestions — before any FactSet call.
- Missing FQL values become typed N/As instead of poisoning a sheet.
- Holdings vs timeseries needed different config models; some formulas need a second-pass fetch.
- Runs only inside FactSet's Programmatic Environment.
Learned
Treat research tooling as a product: validate the config language first, then plan → fetch → derive → export.
Also
- Holdings, timeseries, validate, list, and FDT commands (packaged CLI, v0.3.0)
- AST-gated formula language + topological evaluation (no NetworkX)
- Pydantic configs with cycle checks and typo suggestions
- Rich terminal tables and multi-sheet Excel with hidden helper metrics
What I can say
ML-Ticker was internship work at Merrill Lynch: a Python CLI against FactSet FQL, used on a calculations platform that sits on $500M+ AUM. There is no public GitHub link. I am not going to describe internal tickers, entitlements, or anything that looks like a leak dressed up as a portfolio bullet.
What I will describe is the shape of the system, because that shape is the work.
Sandbox, then graph
FQL is powerful and easy to get wrong. The CLI takes YAML/Pydantic configs (extra fields forbidden), validates a restricted formula language, plans the FactSet calls, then evaluates Python metrics in topological order. Holdings and timeseries are different config models. Some formulas need a second-pass fetch. FDT mode scores a name against peers into Bull / Base / Bear multiples.
It only runs inside FactSet's Programmatic Environment. That is a real constraint, not a footnote.
Why the AST and the DAG
Formulas had to be useful without becoming a code-execution hole: AST denylist, allowlisted builtins, no NetworkX. Interdependent metrics get a DAG, a cycle path, and typo suggestions — before any FactSet call. Missing FQL values become typed N/As instead of poisoning a sheet. Export is Rich tables, multi-sheet Excel with hidden helper metrics, JSON.
The failures that hurt were type failures that looked like data failures. Making the graph refuse a silent coerce was more useful than a prettier command. Treat research tooling as a product: validate the language first.