Python Software Engineer | AWS Data Automation
We’re looking for an experienced Python Software Engineer to join a highly technical data automation team building production-grade, event-driven solutions on AWS..
You’ll help build real-time data extraction services that process long-form, unstructured information, primarily PDF documents.
These services use AI-assisted extraction to identify relevant information, validate and transform it into semi-structured JSON, and make it available to a wider platform as quickly and accurately as possible.
The solutions typically include human-in-the-loop validation, meaning the systems must be designed to manage confidence levels, exceptions and manual review without creating unnecessary delays.
You’ll be working across areas including:
- Building production-grade Python services and applications
- Developing AWS Lambda functions for document ingestion, processing and transformation
- Designing event-driven workflows using Amazon EventBridge
- Processing long-form PDFs and other unstructured data sources
- Converting extracted information into consistent, usable JSON objects
- Integrating AI-assisted extraction into production workflows
- Implementing error handling, retries, monitoring and operational controls
- Supporting human review and validation processes
- Writing clean, tested and maintainable code
- Working closely with other experienced engineers in a highly technical delivery team
Experience
- Strong commercial Python software engineering experience
- Hands-on experience building production workloads with AWS Lambda
- Strong knowledge of Amazon EventBridge
- Experience designing serverless and event-driven architectures
- Understanding of software engineering principles, including testing, maintainability and observability
- Experience integrating services through APIs, events and structured data contracts
- Ability to explain the technical decisions and trade-offs behind the solutions you have built
Desirable
- AI or LLM-assisted document extraction
- Processing PDFs or other unstructured document formats
- OCR, document parsing, chunking or classification
- Building evaluation systems to measure extraction accuracy and prevent regressions
- Confidence scoring and human-in-the-loop workflows
- Terraform or other Infrastructure as Code tooling
- Snowflake
- Schema validation and data-quality controls
- AWS monitoring and observability tooling
If you’ve built production-grade Python applications using Lambda and EventBridge, particularly involving unstructured documents or AI-assisted extraction, I’d love to hear from you.