Python Software Engineer
Opis stanowiska
Job Description
CPGvision is a recognized leader in Trade Promotion Management (TPM), Trade Promotion Optimization (TPO), and Revenue Growth Management (RGM). Leveraging the power of the Salesforce platform, we enable consumer goods companies to achieve their RGM objectives through our fully integrated, user-friendly solution suite.
About the Team
Our Data Science team builds scalable machine learning models for the CPG and Retail industries, helping global brands optimize pricing and promotional strategies. We are looking for a Python Software Engineer - first and foremost a strong software craftsman. Your work will span ML systems and data acquisition pipelines, but what matters most is your ability to write excellent, production-grade Python.
Key Responsibilities
- Software Engineering: Design and develop clean, well-tested, maintainable Python code - from shared internal libraries to production services.
- ML Pipeline Engineering: Build end-to-end ML pipelines (ingestion, feature engineering, training, serving) with gradient boosting models (LightGBM, XGBoost), focusing on reproducibility, performance, and automation.
- Data Acquisition: Develop Python-based scrapers for selected web portals and integrate them with OCR tools to extract structured data.
- Cloud Infrastructure: Deploy and operate workloads on AWS (S3, EC2, Lambda), including containerized jobs, scheduling, and monitoring.
- Collaboration: Partner with Data Scientists to productionize their work, conduct code reviews, and raise the engineering bar across the team.
- Excellent Python skills - the core requirement of this role. We expect a deep understanding of the language internals, not just fluency in using it. Candidates will be assessed primarily on software engineering ability.
- A degree in Computer Science, Data Science, Engineering, Applied Mathematics, or a related field.
- Experience with web scraping (e.g. Scrapy, BeautifulSoup, Selenium, Playwright), including dynamic content and rate limiting.
- Practical experience integrating OCR tools into data pipelines (e.g. Tesseract, cloud-based OCR services).
- Working knowledge of gradient boosting models (LightGBM, XGBoost).
- Working knowledge of AWS (S3, EC2, Lambda).
- Proficiency in Docker.
- Experience with Git and CI/CD workflows.
- Comfortable working in a Linux terminal environment.
- Communicative English.
- Experience with deep learning models for time-series (e.g. Temporal Fusion Transformers).
- Experience with experiment tracking and model registry tools (e.g. MLflow).
- Experience with workflow orchestration (e.g. Airflow, Prefect, Step Functions).
- Familiarity with hyperparameter optimization frameworks (e.g. Optuna).
- Knowledge of Infrastructure as Code (e.g. Terraform, CloudFormation).
- Familiarity with CPG/FMCG or Retail data domains.
- Choice of employment contract or B2B.
- Fully remote work with quarterly on-site meetups (2-3 days).
- Work with large-scale datasets and real impact on decisions of major corporations.
- Access to cloud infrastructure (AWS) and a modern technology stack.
- Opportunity to shape the engineering culture and tooling within the team.
- Mentorship support, Multisport card, and private medical care.
Interesuje Cię ta oferta?