AI Software Engineer — Python, Document Intelligence & AWS (Remote)
nextgen coding companyNew York, NY
$30/hourAPPLY NOW
AI Software Engineer — Python, Document Intelligence & AWS (Remote)
nextgen coding companyNew York, NY
2 days ago
$30/hour
APPLY NOWAbout the role
AI Software Engineer — Python, Document Intelligence & AWSNextGen Coding Company
Location: New York City — Hybrid / In-Person Required
Compensation: $30/hour
Hours: 40 Hours/Week
Engagement: Contract with potential for long-term engagement Eligibility: Must be a U.S. citizen and able to complete a W-9. No visa sponsorship is available for the role.
Role Overview:
NextGen Coding Company is hiring an AI Software Engineer in New York City to join an active enterprise AI project focused on document intelligence, large-scale data processing, and evidence-backed AI analysis. The work is hands-on and engineering-heavy. You will help build a cloud platform that ingests thousands of pages of PDFs, scans, images, HTML, and other data; performs OCR and Python-based processing; structures and indexes the resulting information; and makes the data searchable and usable by modern LLMs. We are looking for a strong builder, not someone whose AI experience consists primarily of calling an LLM API.You should be comfortable jumping directly into an existing project, understanding the architecture, debugging difficult data-processing problems, and shipping production code. What You’ll Work OnBuild production systems primarily in PythonBuild and improve large-scale document ingestion pipelines Process PDFs, scans, images, HTML, and other file formats Build OCR and document-intelligence workflows Classify incoming documents and route them through appropriate processing pipelinesBuild Python post-processing to clean, normalize, validate, and structure extracted information Solve large-file processing issues using chunking, queues, parallel processing, retries, and recovery Extract text, tables, entities, metadata, relationships, and structured records Build ETL and asynchronous data-processing pipelinesBuild hybrid search, vector search, embeddings, RAG, and reranking Integrate Claude, OpenAI, Gemini, Qwen, and other LLMsBuild evidence and citation systems connecting AI outputs to original source material Build and maintain production infrastructure in AWSWork with PostgreSQL, OpenSearch, S3, queues, caching, and APIsBuild integrations including authentication, webhooks, Stripe, and third-party services Write tests for OCR, extraction, retrieval, data processing, APIs, and AI outputs Diagnose production failures and improve system reliability and performance Core Technical SkillsPython: FastAPI, data processing, ETL, APIs, asynchronous/background jobs Document Intelligence: OCR, PDF parsing, scanned documents, OpenCV, PyMuPDF, PaddleOCR, Tesseract, Docling, Unstructured, or similar toolsAI: Claude, OpenAI, Gemini, Qwen, RAG, embeddings, structured outputs, reranking, model orchestration Data: PostgreSQL, SQL, OpenSearch/Elasticsearch, vector databases, RedisAWS: S3, RDS, OpenSearch, Bedrock, EC2/ECS/EKS, Lambda, SQS, IAM, CloudWatchInfrastructure: Docker, Linux, CI/CD, production cloud deployments Frontend experience with React, Next.js, and TypeScript is helpful but is not the primary focus. Who We WantWe want someone who can be given a difficult engineering problem and figure it out. For example:“We have several thousand pages across PDFs, scans, images, and other file formats. Large files are processing inconsistently. Build a reliable AWS pipeline that classifies the files, performs OCR where required, cleans and structures the output with Python, handles failures, indexes the information, and makes the evidence usable by an LLM.”You should be able to break a problem like the one above into an architecture and then actually build it. Strong candidates will have experience building real production systems involving Python, data pipelines, OCR/document processing, AWS, or AI.
Requirements:
- Must be based in New York CityMust be a U.S. citizen
- Must be available approximately 40 hours per week
- Strong Python engineering ability
- Production AWS experience
- Experience with OCR, document processing, data engineering, or similar high-volume processing systems
- Experience integrating modern LLMsStrong backend/API/database fundamentals
- Comfortable independently debugging complex engineering problems
- Able to meet with our team in person in NYCAble to complete a technical engineering screen
Interview Process:
We are intentionally keeping the process straightforward:1. Application Review — Resume, LinkedIn, GitHub, and/or examples of systems you have built.2. Technical Screen — Python, OCR/document processing, AWS, data architecture, and LLM engineering.3. In-Person NYC Meeting — Meet the team, walk through the project, and discuss how you would approach real engineering problems from the platform. We are looking for someone who can join quickly, take ownership, and immediately contribute to a technically ambitious production AI system. About NextGen Coding Company NextGen Coding Company is a U.S.-based software engineering firm building custom software, AI systems, automation platforms, data infrastructure, and enterprise applications. Our engineers work on real production systems across AI/ML, document intelligence, data engineering, financial and compliance technology, and cloud infrastructure. To ApplyPlease send: Resume or LinkedInGitHub and/or portfolio, if available A short description of the most technically difficult production system you have built Any relevant experience with Python, OCR/document processing, AWS, and LLMsNYC candidates only.
Matching similar jobs
JOB OVERVIEW
Salary
$30/hour
Experience level
Senior
Location
New York, NY
Occupation
Software Developers
Industry
Software Publishers
Posted
2 days ago
Tired of running searches?
Rank the roles you'd take once, and matches like these arrive on their own.
CREATE PROFILE