Welcome to DWiM Discussions! #1
Replies: 6 comments 105 replies
|
Hi @pradeeban, My name is Suryansh Maurya, and I'm really excited about the DWiM project. The goal of building a MATLAB-native alternative to Niffler is a fantastic idea, especially with its potential for GSoC 2026. I saw the repository is brand new, and I'd love to jump in and be the first contributor. My programming background is in C++, Python and Java, so I'm very comfortable with data structures, OOP concepts, and building modular systems. I'm keen to apply those principles to help build this framework in MATLAB. After reading the project description, I've outlined a potential phased plan that I think could work well. I'd love to get your thoughts on it. My idea is to build this in a modular way, so each part (retrieval, processing, analysis) can be built and tested separately. Phase 1: Environment Setup & Core Retrieval
Phase 2: Building the Core Processing Pipeline (Niffler Feature Parity)
Phase 3: Integration & The "Scientific Contribution"
That single-environment, end-to-end workflow would be a huge contribution. To get the ball rolling, I'd be happy to tackle the first part of Phase 1. I can:
Let me know what you think of this approach! I'm really looking forward to hearing your feedback and am ready to get started. Best, |
|
Heyy @pradeeban , As I finalize my GSoC proposal for DWiM. However, I wanted to discuss the alternative plans for three specific architectural risks to make sure we are ready for everything. Here is what I currently have planned in the proposal, and the alternative approaches I'd love your feedback on:
What is currently in my proposal: I plan to build the background listener using MATLAB's native What I need your feedback on (Not in the proposal) Before I make this my alternative approach, I want to ask about the ecosystem impact. Are there specific security or cross-platform deployment concerns with shipping compiled DCMTK binaries alongside the
What is currently in my proposal: For the What I need your feedback on (Not in the proposal): Because NIfTI
What is currently in my proposal: Currently, the pipeline flattens complex nested DICOM sequences in memory for extremely fast writing to tables/CSVs. If this causes RAM exhaustion on petabyte scale datasets, my alternative is to move to MATLAB's out of core data handling ( What I need your feedback on (Not in the proposal): From your preceptive , do you prefer keeping DWiM strictly reliant on in memory operations for architectural simplicity, or does moving toward SQLite/out of core datastores for the batch exporter align better with the project's long-term vision? I wanted to add these in my proposal but i wanted to have your view on this first. Plus my proposal is already 37 pages long and it might get more than 40 pages if add these alternatives. will that be okay ? |
|
Just adding a small thought here — earlier @pradeeban mentioned that Niffler itself was the result of several years of work and contributions from multiple developers, so it’s completely reasonable that this project focuses on building a solid foundation first instead of trying to reproduce everything within a single GSoC period. Even though development is faster now with modern tools and MATLAB Live’s native AI integration, he also suggested keeping the 350‑hour GSoC scope realistic and clearly separating definite goals from “if time permits / optional” items. That might help keep expectations clear and make it easier for future contributors to continue extending the project. Your work here is excellent and the way you're thinking about the long‑term direction of the project is really valuable. |
|
Thanks a lot for your input @KrishanYadav333. To clarify, most of the groundwork for this project has already been completed through my previous PRs, and @pradeeban is already aware of the overall direction and has approved the approach. And regarding the 350-hour scope, I actually raised a similar concern earlier in this discussion: #7 (comment), where I suggested your approach might not fit in the 350 hour timeline. :) From the very beginning, I’ve been careful to respect the 350-hour GSoC constraint as my entire progress from day 1 is planned according to that only. At the same time, since the major foundational work is already in place, I planned to dedicate more effort to the later phases, which I believe are the most critical parts of the project. And according to me if we already limit our developer mindset without giving in extra efforts in the provided time than it's not a good habit. By keeping the quality of our efforts on top and doing our best we can achieve great results during the provided timeframe. That was all that I meant in my previous post. Ultimately, I completely agree that clear communication and well defined scope are important for a project of this scale. Again since this is a completely new project, so non of the approaches are right or wrong as stated by @pradeeban. And I respect and appreciate your efforts to this repository. Please keep it up! |
|
@pradeeban I have submitted my proposal in the portal and will be updating it with fine refinements and updates before the deadline. Please have a look and if you feel anything that can be improved, I'd be glad to implement that. My other two proposals are in the work and will share that with you in 1 or 2 days for the review. Again Thankyou for everything :D |
|
Heyy @pradeeban, I hope you’re doing great. I wanted to share something that has been quite meaningful to me personally. Over the past few months, I had the chance to work on a research paper titled “TabFlowM: Lightweight Flow Matching for Mixed-Type Tabular Data Synthesis in Latent Space”, which is currently under review at TMLR (OpenReview). This is during my intern at NTU Singapore, IIT BHU and the University of Warwick, and honestly, the experience pushed me to learn a lot and i have grown beyond what I initially thought I could handle. Here’s the link: https://openreview.net/forum?id=t5kygrpSIz&referrer=%5BAuthor%20Console%5D(%2Fgroup%3Fid%3DTMLR%2FAuthors%23your-submissions) I’m sharing this not just to showcase my work, but because I genuinely value your guidance and perspective. If you get some time to read through this, even briefly, your feedback would mean a lot to me. And I know you are really busy so if you don't get the time it's completely fine :) Thank you so much for your time, it really means a lot. |
Uh oh!
There was an error while loading. Please reload this page.
👋 Welcome!
We’re using Discussions as a place to connect with other members of our community. We hope that you:
build together 💪.
To get started, comment below with an introduction of yourself and tell us about what you do with this community.
All reactions