Lead Data Engineer at Convex Insurance (2023-12 – Present)
Developed the first near-real-time pipeline for third-party data sources, such as Verisk insurance, enhanced trust and integration with one of the company's key departments within the data platform.
Reducing
Fivetran costs for ingestion, and deprecating Dataiku, a duplicate tool used only by certain business areas. I implemented centralized observability for our toolset using Datadog, improving time to resolution by 30%. Later, as a Lead Data Engineer, I'm guiding and implementing the transition from a reinsurance legacy system to a more modern provider. Designing data models, pipelines, and architecture for integrating the new system within the company, working with business stakeholders, enterprise architects, training testers, and other engineers.
- Designed ingestion and reverse ETL architecture for a reinsurance system using Prefect, Boomi, Snowflake, dbt, and Python.
- Training of the company's testers who were not familiar with data projects in general. Set up validation automations using Streamlit and dbt packages.
- Infrastructure design and deployment using Terraform, new environments creation for validation with the provider went from 1 week to less than a day.
- Active participation in planning for the project timeline and budgeting, definition of tooling, stakeholder management and training of the team.
- Defined go-live, sunsetting and transition of as-is toolset to new system with Finance, Actuarial and Ceded Operation team.
- Defined and implemented SOX requirements strategy for external auditing
Senior Data Engineer at Convex Insurance (2023-12 – 2023-12)
- Reduced Fivetran spending by 85% using Prefect as an external orchestrator only syncing data once it was completed in the source, avoiding unnecessary MAR charges.
- Migration off the company monorepo to a new structured data-engineering repository, re-structuring CI/CD pipelines, and documentation without losing history and affecting other engineers' ongoing work.
- Designed and implemented the first near-real-time pipeline of the company, integrating the portfolio optimization team which was the only remaining business area not fully using the data platform in the company.
- Training and mentoring of one returnee and 2 other junior members on the company stack. Restructure of the onboarding, reducing time to first contribution to the data project from 4 weeks to 2.5 weeks.
- Designed and implemented the first centralized monitoring of the data team, using Datadog to track ingestion (Fivetran), transformation (dbt and Scala), and orchestration (Prefect). Improving observability for engineers and stakeholders.
Lead Data Engineer at Inventa (2021-05 – 2023-12)
First data hire of the company, initially as a data scientist. I led all analytics practice as the only data team member until reaching a Series B investment round, implementing all the ETL, machine learning models, experimentation, and dashboarding, with direct participation in product and strategic decisions. Later, with the addition of team members, managed and mentored junior employees, and led the migration from the initial stack to a more robust implementation of the modern data stack.
- Migration of data stacks to the Modern Data Stack standard, with no data downtime to stakeholders.
- Designed a Snowflake data warehouse using dbt, training and leading 9 data scientists to contribute to the project.
- Snowflake cost optimization reducing expenses by 50% by tuning warehouse sizes, parameters, and materializations in dbt.
- Lead employee brand awareness initiatives, creating content with Hex and dbt, such as blog posts, case studies, and presentations.
- Designing of hiring process, defining profiles, training for HR recruiters and technical interviews, and being directly responsible for the hiring of 70% of the company data team.
- Led production data migrations after changing the product stack, developed a process, and migrated the company users, orders, and payments with more than 98% success.
- Designed most of the data platform processes, such as on-call, bug triage, documentation, and code standards. Reducing time to resolution on bugs by 45%.
Data Scientist at Inventa (2021-05 – 2023-12)
- Setup of all the ingestion of production data using S3 as a data lake, SNS, SQS, Lambda, RDS, and Airflow for ETL.
- Modeling and implementation of the data warehouse, managing the infrastructure with access control and data quality tests.
- Deploy of lead scoring model used by the inside sales department that increased pickup rates and GMV per sales agent by 115% and 49% respectively.
- Product analytics and experimentation alongside the product team, being responsible for more than 25 experiments that helped the company to address key changes in the platform.
- Direct sourcing of metrics and analysis for the investment round alongside the founders, creating reliable and audited data reports that helped the company raise 2 rounds of investment.
- Orchestration and development of several scrapers for tracking order deliveries using AWS services such as SQS, Airflow, and Lambda Functions. Auto tracking of 85% of the packages without having a paid integration with transport companies.
Gas and Power Research/Data Analyst at Wood Mackenzie (2019-02 – 2021-05)
Wood Mackenzie team had a very non-automated and manual workflow for market analysis and report creation. I was able to fully automate the company data flows such as public data retrieval, and model forecasts to deliver our product to stakeholders. Increasing the number of assets delivered per year from 6 to 14, with new products such as live dashboards.
- Creation of a full public data scraping solution, orchestrating the extraction of energy sector data from 6 countries.
- Setup and management of a data warehouse using SQL Server, creating transformations for data analysis and for reporting automation generating company standard reports and PowerBI dashboard for external clients.
- Modeling and presentation of energy market insights for existing and prospective clients.
Metadata QC Archivist at Olympic Channel (2017-01 – 2019-02)
The main challenge of the Olympic Channel archiving service was to have consistent quality control over the metadata for thousands of hours of content generated in each Olympics. I supported the team by creating automation tools and processes that reduced the time for fulfilling content requests by more than 85%. I also restructured their entire training program for freelancers, from 5 full days of training to 3 days part-time, without losing the quality of hired candidates.
- Training of more than 100 local Spanish graduates to create a qualified freelancer workforce for high-demand operations.
- Coordination of the video logging operation in 2 Olympics, being responsible for managing a 20 people team.
- Providing technical solutions such as scripting for task automation (Windows and Linux), manipulation, and visualization of metadata stored as XML files using C++, Python, and PHP as programming languages.
System Developer Researcher at LIOC COPPE - UFRJ (2013-04 – 2016-12)
I've participated in several R&D projects from the Oceanography Instrumentation Laboratory from COPPE - UFRJ as a researcher intern and later as a hired developer. During the period in the laboratory, I was one of the 3 undergraduates developing extremely low-power embedded systems, that were used to power an entirely produced autonomous glider, an industrial meteorological buoy, and tide gauges.
- Development of C++ drivers to interact with industrial sensors and an Iridium Satellite module.
- Planning and development of the core system for an embedded system with power limitations. Using a Raspberry Pi, Texas MSP430, and BeagleBone Black.
- Development of Linux scripts for performing system tasks such as backups, sending and reading emails, and manipulating the IO pins of the boards.