Principal AI/ML Engineer, Semantic Data
major league soccerBrooklyn, NY
yesterday
Occupations
Data ScientistsDatabase ArchitectsData Warehousing SpecialistsIndustries
Computer Systems Design ServicesCustom Computer Programming ServicesSoftware PublishersAbout the role
Overview
In this role you will design and scale the semantic data layer powering MLS fan intelligence and AI-driven decisioning. You’ll build knowledge graphs, retrieval systems, and RAG-enabled AI capabilities that unify fan, content, and business data. You will work across ML, data engineering, and product teams to deliver production-grade AI infrastructure with strong grounding and explainability. This is a hands-on, systems-focused position that shapes how MLS reasons about data and AI at scale. You’ll join a mission-driven team that values collaboration, performance, and responsible AI.
Compensation / Benefitscomprehensive medical, dental, and vision coverage$500 wellness reimbursement Holiday and PTOcareer development and ongoing educationon-the-job training and feedbackin-person collaboration with flexible remote day(s)
Responsibilities Design and implement embedding pipelines for fan data, content, metadata and behavioral signals Build metadata and enrichment systems to normalize and structure enterprise data for AI use Develop knowledge bases and retrieval systems using vector databases and hybrid search Create context assembly pipelines combining structured data, documents, APIs, and historical outputs Enable AI systems to operate on unified semantic representations rather than raw data Architect and manage knowledge graphs representing fan, content, and business entity relationships Define and maintain a semantic layer standardizing metrics, features, and business concepts Design ontologies, taxonomies, and entity models for fan behavior and identity Implement graph-based reasoning and enrichment workflows Ensure semantic consistency across analytics, ML, and operational systems Design and build retrieval-augmented generation systems grounded in semantic data Integrate LLMs for reasoning over structured and unstructured data Develop pipelines translating natural language into structured outputs such as queries and analytical tasks Build and optimize context pipelines improving LLM grounding and factual accuracy Evaluate and integrate open-weight models for domain-specific reasoning Fine-tune or adapt models using parameter-efficient techniques Support deployment of LLM systems in private or on-prem GPU environments Optimize inference workflows for latency, cost, and scalability Enable LLM-driven workflows that reason over semantic data and retrieval systems Build scalable, production-grade services and APIs for semantic and AI systems Work with vector and graph databases to support retrieval and reasoning Integrate structured data, documents, APIs, and model outputs Partner with data engineering on batch and real-time pipelines Ensure systems meet performance and reliability requirements Design evaluation frameworks for retrieval quality and LLM output correctness Monitor system performance, relevance, and model behavior Establish guardrails for explainability, traceability, and data attribution Ensure safe and reliable generation of structured outputs Mitigate risks related to bias, data leakage, and inconsistencies Collaborate with product, analytics, and engineering teams on AI use cases Translate business problems into systems combining semantic data and LLM reasoning Partner with ML teams to improve model performance through better grounding Mentor engineers and establish best practices
Key requirements 8–10+ years of experience in ML engineering, data systems, or applied AIStrong expertise in Python, SQL, and production software engineering Deep experience with semantic data modeling, ontologies, and entity resolution Hands-on experience with embeddings, vector search, and retrieval systems Experience building and deploying LLM-powered systems including RAGExperience building production-grade AI systems at scale Strong understanding of distributed systems and data architecturecollaboration across product, analytics and engineeringmentorship and best-practices sharingstrong communication of complex conceptssemantic data modelingontologies and entity resolutionembeddings
Matching similar jobs
JOB OVERVIEW
Experience level
Lead
Location
Brooklyn, NY
Occupation
Data Scientists
Industry
Computer Systems Design Services
Posted
yesterday
Tired of running searches?
Rank the roles you'd take once, and matches like these arrive on their own.
CREATE PROFILE