Machine Learning and Software Engineer
Email / LinkedIn / GitHub / Scholar
(September 2024 - Now): Machine Learning Engineer - Moonvalley (https://www.moonvalley.com/)
- Responsible for preparing the next big thing in text-to-video and text-to-image models.
(March 2024 – September 2024): Machine Learning Consultant - Equinox AI (https://equinoxai.com/)
- Responsible for optimizing pipelines for high-volume real-time algorithmic trading systems, using a mixture of big-data time-series techniques
and reinforcement learning trading agents.
(November 2018 – February 2021): Lead Machine Learning Engineer - Deepsea Technologies (https://deepsea.ai/)
As a Lead, I was responsible for the following projects:
-
Uncertainty Estimation/ Anomaly Detection.
Used Deep Quantile Networks to forecast and report anomalies on tankers and cargo ships of 5 major worldwide corporations. Proposed better engine usage schemas, reduced fuel costs by 4%, and detected oil theft on various ships. -
Deep Learning Based Ship Route Optimization.
Used Generative Adversarial Networks to Simulate Ship Performance under different weather conditions and created a better Deep Learning boosted variant of the Graph-Based A*(A star) algorithm. Managed to automatically create routes for over 100 ships in real-time, with the algorithm outperforming current competitor solutions by 8-10%. -
AIS System Based - Fuel and Hull Fouling Estimation over Sparse Data. Used Time Series Deep Neural Networks and Feature Engineering based on Geolocation Trajectories to report fuel and power consumption based solely on sparse and corrupt AIS Data. Created a robust system that can classify performance deterioration and assess the overall state of a ship.
(Dec 2017 – Nov 2018): Machine Learning Engineer - Deepsea Technologies (https://deepsea.ai/)
As an ML Engineer, I was responsible for the following projects:
-
Product Pipeline Design.
Successfully containerized and published all Deep Learning Projects with Docker and Flask. Created a multi-worker high-availability system with Gunicorn and balanced incoming traffic using Nginx. Finally, all projects were deployed on AWS Cloud, and a combination of Jenkins and Coverage was used to maintain CI/CD. That led to the creation of a “click-and-play” Deep-Learning project base for over 200 ships and five different companies. -
Modeling “Aging” Phenomenon on Mechanical Parts.
In this project, we aimed to model the magnitude of “aging” on different mechanical parts and automatically suggest their Remaining Useful Life (RUL). We used a combination of classical Machine Learning and Dilated Convolutional Neural Networks and built a real-time monitor/tracker deployed on every ship. We reduced ship maintenance costs by 10% by suggesting maintenance and repair.
|
Programming Languages
CI / CD |
Machine Learning
Data Storage and Analytics |
(Oct 2021 – Jan 2025): PhD Candidate, University of Warsaw
Research Area: Reasoning in Deep Learning Models
My research focuses on exploring the reasoning capabilities of deep learning models in the challenging modalities of vision and language and their interplay.
- My initial project, Measuring CLEVRness: Black-box Testing of Visual Reasoning Models, showcased at ICLR 2022, aimed to advance our grasp of reasoning in visual-language scenarios through a novel dual-agent setup. This method, diverging from standard visual QA and adversarial tests, bypasses the need for direct model access like gradients or probabilities. Instead, inspired by zero-sum game theory, it transforms a visual question-answering scenario into a contest between a reasoning agent and a scene-manipulating agent, evaluating reasoning via significant, context-relevant scene changes. Thus, we unveiled the shortcomings of prevalent data-driven approaches, challenging the accuracy of their claimed super-human visual QA performances. We also uncovered model biases masquerading as logical reasoning by analyzing agent interactions. By leveraging these results, we introduced fine-tuning strategies, leading to more robust multimodal reasoning akin to human processes.
- In the subsequent work, A Simple, Yet Effective Approach to Finding Biases in Code Generation presented at ACL 2023, we explored the reasoning capabilities of various Large Language Models within the evolving and critical field of code generation. Our work breaks new ground by formalizing and evidencing three primary bias sources in coding challenges, termed Blocks of Influence: function names, problem specifications, and examples. A novel framework designed to automate the assessment of LLMs was introduced, revealing issues related to models' over-reliance on example demonstrations, biases in naming conventions, and difficulties in linking language descriptions to corresponding tasks. We validated our claims by investigating several open-source models by pinpointing the exact sources of memorized code and naming dependence in their datasets. Building on these insights, we utilized these identified biases as training augmentations against various open-source LLMs. We achieved significant robustness against such phenomena, reflected in improved coding abilities across all tested models.
- In my latest study, Beyond Lines and Circles: Unveiling the Geometric Reasoning Gap in Large Language Models, we delved into the intricate world of constructive geometry with LLMs, introducing three key methodological innovations. Our first innovation is a dynamic prompting technique inspired by Retrieval Augmented Generation, which significantly improves LLMs' adaptability and contextual awareness. Based on their relevance in solving progressively complex problems, this technique uses past interactions and reranks results. Secondly, we introduced a novel prompting technique that effectively tackles LLMs' need for 2D spatial awareness, facilitating precise problem extraction and applicability to models without inherent visual capabilities. Finally, the centerpiece of our research is the novel multi-domain and multi-role agent framework. Our approach demonstrates significant improvements in performance on mathematical and geometric QA tasks, outperforming vanilla multi-agent systems and setting new benchmarks for the complex problem-solving capabilities of LLMs. Crucially, our methodologies prove highly transferable to smaller LLMs, which greatly benefit from the enhanced reasoning and problem-solving capacities introduced by our innovations.
(Oct 2021 – Now): Teaching Assistant, University of Warsaw
Responsible for conducting laboratories in the Natural Language Processing
and Visual Recognition courses for the master’s degree study cycle at the University of Warsaw.
Proud to have supervised the creation of the student project: Distilled HerBert (https://huggingface.co/BartekK/distilHerBERT-base-cased) a BERT-based Language Model trained on Polish Language.
(March 2021 – Oct 2021): Research Assistant, University of Warsaw
Research Area: Visual Reasoning and Neurosymbolic Approaches.
Under the supervision of Henryk Michalewski and Mateusz Malinowski,
we investigate the effect of Neurosymbolic Neural Networks as well as
Adversarial Agents towards testing the limits of VR/VQA tasks.
During this work, an interactive application that acts as a neural
network testbed has been created, enabling a better understanding of
pretrained visual reasoning architectures.
Beyond Lines and Circles: Unveiling the Geometric Reasoning Gap in Large Language Models (EMNLP 2024) PDF ▶ Play
A Simple, Yet Effective Approach to Finding Biases in Code Generation (ACL 2023) PDF ▶ Play
Measuring CLEVRness: Blackbox testing of Visual Reasoning Models (ICLR 2022) PDF ▶ Play
MAIN: Multi-Head Attention Imputation Networks (IJCNN 2021) PDF ▶ Play
Optimized Generation of Hardware CNN Inference Engines (MDPI Technologies 2020) PDF ▶ Play
TF2FPGA: A Framework for Projecting and Accelerating CNNs on FPGA Platforms (IEEE Mocast 2019)
PDF
▶ Play
|
PhD Studies Topic: Reasoning in Deep Learning Models
Master of Science in Data Science and Machine Learning
Summer School - Graph Theory and Randomized Computing
Bachelor and Master of Engineering in Electrical and Computer Engineering |
Worked and published under the OPUS-15 Grant with Prof.Henryk Michalewski
Greek: Native |