Abstract
In many ML use cases, model performance is highly dependent on the quality of the features they are trained and inference on. One of the important dimensions of feature quality is the freshness of the data. Therefore, it is critical to ensure that the features remain up-to-date to the problem being solved.
The presentation will cover the impact of feature freshness on model performance based on experiments both in training data and inference data. We will also discuss various strategies and techniques that can be used to improve feature freshness, including in streaming and batch feature processing. It will also discuss the challenges and tradeoffs that come with implementing these strategies in large scale machine learning systems, such as the computational cost and scalability issues.
By keeping the features fresh and relevant, organizations can achieve better results and stay ahead of the competition in today's rapidly evolving data-driven landscape.
Interview
My current area of focus revolves around developing techniques to prepare data for machine learning inference on a large scale. At the same time, I aim to enhance reliability, improve efficiency, and minimize latency in the process.
I would like to share our learnings while working on these projects with the industry.
The target audience would be experienced technologists in the industry who work on large scale data processing for machine learning.
There are a few key takeaways:
- Improving data freshness is becoming more and more important in ML tasks
- However not all your data need to be super fresh. Optimize for ROI instead of freshness alone
- Design your system end to end, instead of focusing on localized optimization
Topics
QCon New York 2023 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
Part of the track
MLOps: Navigating the Terrain of Large-Scale Models Hosted by Bozhao (Bo) Yu Founder @BentoML.aiFrom the same track
Wednesday 14 June
10:35 Salon D Session ML Infrastructure Introducing the Hendrix ML Platform: An Evolution of Spotify’s ML Infrastructure Divita Vohra, Mike Seid The rapid advancement of artificial intelligence and machine learning technology has led to exponential growth in the open-source ML ecosystem. 11:50 Salon D Session Machine Learning Improve Feature Freshness in Large Scale ML Data Processing Zhongliang Liang Engineering Manager @Facebook AI Infra In many ML use cases, model performance is highly dependent on the quality of the features they are trained and inference on. One of the important dimensions of feature quality is the freshness of the data. 13:40 Carroll Gardens Unconference Unconference: MLOps What is an unconference? An unconference is a participant-driven meeting. Attendees come together, bringing their challenges and relying on the experience and know-how of their peers for solutions. 14:55 Salon D Session AI/ML A Bicycle for the (AI) Mind: GPT-4 + Tools Sherwin Wu, Atty Eleti OpenAI recently introduced GPT-3.5 Turbo and GPT-4, the latest in its series of language models that also power ChatGPT. 16:10 Salon D Session MLOps Platform and Features MLEs, a Scalable and Product-Centric Approach for High Performing Data Products Massimo Belloni Data Science Manager @Bumble In this talk, we would go through the lessons learnt in the last couple of years around organising a Data Science Team and the Machine Learning Engineering efforts at Bumble Inc. 17:25 Salon D Panel Panel: Navigating the Future: LLM in Production Sherwin Wu, Hien Luu, Rishab Ramanathan Our panel is a conversation that aim to explore the practical and operational challenges of implementing LLMs in production. Each of our panelists will share their experiences and insights within their respective organizations.