
In Progress
Posted
CUDA Developer Needed – Long-Term Remote Work (AnyDesk/Remote Access) I am looking for an experienced CUDA/C++ developer to help me complete GPU optimization tasks. This is a long-term collaboration opportunity for someone with strong experience in CUDA kernels, PyTorch, and performance optimization. Work setup: Remote work through AnyDesk or another remote desktop tool. You will work directly on my development environment. Tasks involve debugging, improving, and optimizing CUDA implementations. Each task typically requires around 3 hours of focused work. Payment: Rate: $8–$10 per hour depending on experience. Payment will be released after 2 weeks of completed work. This is intended to become a long-term working relationship with regular tasks. Required skills: Strong C++ programming skills. CUDA kernel development experience. Understanding of GPU memory optimization, CUDA streams, and kernel performance. Experience debugging CUDA correctness issues. Knowledge of PyTorch CUDA extensions is a plus. Experience with Linux, Git, and benchmarking tools. Current type of tasks: Fix CUDA kernel correctness issues. Compare CUDA implementations against PyTorch reference implementations. Debug tensor indexing, convolution operations, normalization layers, and residual connections. Optimize GPU performance after correctness is achieved. To apply, please provide: 1. Your CUDA/C++ experience. 2. Examples of CUDA projects you have completed. 3. Your hourly availability. 4. Your preferred remote access method. This can become a long-term position for the right developer.
Project ID: 40610942
43 proposals
Remote project
Active 3 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs

Hi! I have 3.6 years of professional experience in C++ with expertise in Qt, MFC, Win32 API, multithreading, socket programming, and performance optimization. While I’m currently expanding my CUDA skills, I have a strong debugging mindset and can quickly understand complex codebases. I’m available 3–4 hours daily, comfortable working via AnyDesk, and interested in a long-term collaboration. I would be happy to discuss how I can contribute to your project.
$9 USD in 40 days
0.0
0.0
43 freelancers are bidding on average $12 USD/hour for this job

Hi, this reads like targeted GPU debugging work inside an existing environment rather than greenfield development, and I’m used to stepping into live systems to isolate correctness faults before tuning performance. The real engineering risk here is not raw optimization first; it’s locking down numerical correctness across indexing, convolution, normalization, and residual paths so performance changes do not hide invalid outputs. The closest work in my background is Custom Feature Development & Integration, where I dropped into an existing codebase, reviewed the implementation, isolated friction points, and delivered focused fixes safely. I’d also point to Python Bug Localization Using Transformer Models (CodeBERT + TreeBERT) as relevant to structured debugging and fault isolation. I usually approach systems by separating repro, validation, and optimization. For this kind of work, that means first pinning down deterministic failing cases, then tracing where outputs diverge from the reference behavior, then profiling only after correctness is stable. That said, your core need is deep CUDA/C++ kernel specialization, and I should be direct that this is outside my strongest lane. If useful, I can still help review the debugging workflow and validation strategy around the PyTorch reference comparisons. Thanks, Hercules
$50 USD in 40 days
7.2
7.2

As an experienced software development company, our proficiency and expertise in C++ programming, CUDA kernel development, PyTorch, GPU memory optimization, and CUDA streams are highly aligned with your project requirements. We've successfully tackled similar projects involving tensor indexing, convolution operations, normalization layers, residual connections, and subsequent performance optimization. Our adeptness with Linux, Git, and benchmarking tools will ensure smooth communication and efficient collaboration throughout the project. Furthermore, I would like to assure you that we understand the intricacies of Remote Work via desktop tools like AnyDesk as we have a proven track record in remote collaborations. Offering our services at an hourly rate of $8–$10 depending on experience, we guarantee that the task completion will be accurate and punctual - usually within three hours of focused effort per task. In conclusion, what sets us apart is our distinctive fusion of technical excellence and customer-centric approach. We're not just builders - we're problem solvers who prioritize turning ideas into scalable digital products tailored to meet specific client needs. Given a chance, we promise to deliver secure, high-performing solutions with a future focus that will coincide perfectly with your long-term goals. Let's embark on this journey together and build something exceptional!
$10 USD in 40 days
6.8
6.8

Hello I have gone through your specific requirement for CUDA optimization. I would use Nsight Compute before nvprof because it gives kernel level metrics on current CUDA while nvprof is legacy. I will work inside your Linux setup over AnyDesk and debug CUDA kernels against PyTorch CUDA extensions, then tune streams and memory once correctness matches. But first I want repeatable benchmarks. Worked with AI clients on RTX 4090 GPU work and fixed CUDA bottlenecks. Private code, screenshots I can send. Which GPU and CUDA version are you running? Need to check current benchmark flow first. Free for a quick call this week? Dev S.
$15 USD in 40 days
6.7
6.7

Hi, I noticed that you want to boost your CUDA kernels without messing up the current performance. Sometimes just a small tweak in how memory is handled can significantly speed things up; that’s where I find the sweet spot. I can handle debugging, fixing correctness issues, and make sure your kernel runs smoothly and fast. Did you know that optimizing tensor operations like in Labop-Freelancer or CarCRM can really change overall system performance? Let's talk about a plan and build something bigger together. Regards, Nick.
$7 USD in 3 days
7.4
7.4

Your CUDA kernel implementations will fail silently if tensor strides are misaligned during convolution operations. This causes incorrect outputs that look "close enough" during testing but break under production load with varied batch sizes. Quick questions - are you benchmarking against cuDNN baselines to validate correctness before optimizing? And what GPU memory constraints are you working within for your largest model layers? Here is the architectural approach: - CUDA KERNEL OPTIMIZATION: Profile memory access patterns using Nsight Compute to eliminate bank conflicts and coalesced memory issues that kill throughput. - PYTORCH CUDA EXTENSIONS: Build custom autograd functions with proper backward pass implementations that match PyTorch's reference gradient flow exactly. - DEBUGGING CORRECTNESS: Implement tensor validation layers that compare intermediate outputs against PyTorch at every operation to isolate indexing bugs before they propagate. I've debugged similar CUDA correctness issues for 2 ML infrastructure teams where "close enough" outputs caused model divergence after 10K training steps. Let's schedule a 20-minute technical call to review your current kernel implementations and identify the root cause.
$8 USD in 30 days
5.7
5.7

Hello I’m AbdulHamid, a senior developer with extensive experience in CUDA and C++ programming, especially in GPU optimization and kernel development. I’ve successfully debugged and optimized CUDA kernels for deep learning projects, including PyTorch extensions, ensuring correctness and boosting performance through memory optimization and efficient use of CUDA streams. I’m comfortable working remotely via AnyDesk or similar tools and can dedicate over 8 hours daily to meet your deadlines. I’d be glad to review your current CUDA implementations, fix correctness issues, and enhance performance based on rigorous benchmarking. Could you share more about the specific CUDA kernels or modules you’re focusing on right now? Best regards, AbdulHamid
$10 USD in 40 days
5.2
5.2

✋ Hi There!!! ✋ The Goal of the project:- PROVIDE LONG TERM CUDA AND C++ DEVELOPMENT SUPPORT FOR GPU KERNEL DEBUGGING AND PERFORMANCE OPTIMIZATION. <-- Questions --> Are the current PyTorch reference implementations already documented, or will comparison baselines need to be built from scratch during debugging. For remote access, is AnyDesk the fixed requirement, or would another tool work as long as it connects reliably to the development environment. Similar long term CUDA optimization work has been done before, covering kernel correctness debugging, tensor indexing and convolution fixes, normalization and residual connection troubleshooting, GPU memory and stream optimization, and ongoing collaboration through remote desktop access. Looking forward to chat with you for make a deal Best Regards Elisha Mariam!
$9 USD in 40 days
4.9
4.9

Hello, After reviewing your project requirements, I understand that you need an experienced CUDA/C++ developer to debug, optimize, and improve GPU-based implementations through remote collaboration. I have experience with C++, Python, Linux environments, GPU optimization concepts, debugging, and machine learning workflows involving PyTorch. The main challenge in CUDA optimization is balancing kernel correctness and performance while handling memory usage, tensor operations, indexing issues, and GPU execution efficiency. I can help debug CUDA kernels, compare implementations against PyTorch references, optimize bottlenecks, and improve overall GPU performance using proper profiling and benchmarking methods. A couple of quick questions: • What CUDA version, GPU hardware, and development environment are you currently using? • Are the existing CUDA kernels already integrated with PyTorch extensions, or are they standalone implementations? I would be glad to support your long-term GPU optimization work and collaborate directly through your preferred remote access method. Best regards, Carlos
$9 USD in 40 days
4.5
4.5

Hello there, we are a team of senior AI ML Full Stack Web and Mobile App Developers. Please, send me a message to discuss the work and finish in no time. Thanks Ashish Kumar.
$9 USD in 40 days
4.4
4.4

Hi there ! On the correctness bugs you describe, in my experience they are nearly always one of three things: - Input tensor not contiguous, so your stride maths and PyTorch quietly disagree - Missing boundary guard, so the last block reads past the end and only fails on some shapes - fp16 accumulation, where the kernel is actually right and only the tolerance is wrong So I run compute-sanitizer first for out of bounds and races, then a strict allclose against the PyTorch reference. Only once it is correct do I open Nsight Compute. Availability: 4 hours a day, overlapping your hours. Remote access: AnyDesk is fine. If the code can sit in a private git repo, or I get SSH, it is faster for both of us. One ask: fund each task as a milestone rather than paying after two weeks. On that basis I can start today. Timeline : 3 hours per task Budget : $17/hr Thanks!
$17 USD in 40 days
3.6
3.6

Hi, I can support your long-term CUDA/C++ development tasks through remote access, including debugging, correctness validation, PyTorch comparison, benchmarking, and GPU performance optimization. The best solution is to first review each task, current CUDA implementation, expected PyTorch reference output, tensor shapes, test cases, and performance target. I’ll then fix correctness issues, validate outputs, debug indexing/convolution/normalization/residual logic, and optimize kernel performance after the results are stable. I’m comfortable with C++, CUDA kernels, Python, PyTorch CUDA extensions, Linux, Git, GPU memory optimization, CUDA streams, tensor indexing, benchmarking, profiling, and performance tuning. I can help with: * CUDA kernel debugging * Correctness checks against PyTorch * Tensor indexing fixes * Convolution and normalization debugging * Residual connection issue review * GPU memory optimization * CUDA stream usage * Benchmarking and profiling * Linux/Git workflow support * Task-wise progress updates Availability: I can work task-wise and handle 2–3 focused tasks per day depending on complexity. Preferred remote access method: AnyDesk or TeamViewer. I’m comfortable working directly on your development environment through remote connection and can follow technical instructions carefully for long-term collaboration. Best regards Ankit
$10 USD in 40 days
3.7
3.7

Hello, I hope this message finds you well. I understand that you are seeking an experienced CUDA/C++ developer to assist with GPU optimization tasks, and I would be thrilled to collaborate on this long-term project. With over five years of experience in CUDA kernel development and performance optimization, I am well-equipped to tackle the challenges you outlined. My proficiency in C++ and knowledge of GPU memory optimization, along with debugging CUDA correctness issues, positions me to effectively contribute to your project. I have also worked extensively with PyTorch and have a strong understanding of CUDA extensions. To ensure we achieve your goals, I propose the following approach: - Conduct a thorough assessment of existing CUDA implementations to identify and fix correctness issues. - Optimize GPU performance through targeted debugging of tensor indexing, convolution operations, and normalization layers. - Compare CUDA implementations against PyTorch references to ensure accuracy and efficiency. - Utilize benchmarking tools to measure performance improvements post-optimization. I am eager to start working on this project and confident in delivering quality results within the specified deadlines. I am available for remote access via AnyDesk and can start immediately. Let’s discuss any further details to move forward. Thank you for considering my proposal.
$7 USD in 40 days
3.2
3.2

Hi, "CUDA KERNEL DEBUGGING AND GPU OPTIMIZATION" — you need someone who can fix correctness first, then improve performance without breaking PyTorch compatibility. I work comfortably with C++, Python, Linux, Git, and GPU debugging. My approach is to validate every CUDA kernel against the PyTorch reference implementation before optimizing memory access, indexing, streams, and occupancy. That avoids chasing performance on incorrect results and makes each optimization measurable. I'm available for regular remote sessions through AnyDesk or your preferred remote desktop tool, and I'm comfortable working directly in your development environment for focused debugging tasks. Which GPU architecture are you currently targeting, and are the kernels written as native CUDA or PyTorch CUDA extensions? Looking forward to working with you.
$9 USD in 40 days
3.0
3.0

Hello, I just finished reading your brief for "Experienced CUDA/C++ Developer - Remote Long-Term Collaboration -- 3", and I'm genuinely excited about the chance to work on it. I don’t just want to tick the boxes here; my aim is to hand you something that genuinely goes beyond what you pictured. From your brief I can see this involves ai — all areas we handle in-house. We specialise in C Programming, Python, Linux, CUDA, which lines up directly with what you need. How we'd approach it: - Clarify the use-case, inputs and the exact output you expect - Build and integrate the model / automation pipeline - Evaluate accuracy and tune against real examples - Deploy with monitoring and a clear handover Delivering at the scale of 10 per is no problem for us — we're set up for volume without dropping quality. Happy to work hourly with transparent time tracking and regular check-ins. Expect smooth progress updates throughout, full respect for your specifications, and zero surprises along the way. I can start right away and keep you updated at every step — let’s make this a great one. Best regards, FreeLancers360 Let’s connect in chat and get started — message me anytime and I’ll reply right away!
$7 USD in 5 days
3.3
3.3

hi,there.. your project requires more than writing cuda code because correctness must come before optimization. i can help debug kernel logic, validate outputs against pytorch reference implementations, resolve tensor indexing and memory issues, and then optimize performance using shared memory, memory access patterns, streams, and profiling tools. i am comfortable working in linux through remote access, documenting every change, and collaborating on recurring optimization tasks with a structured debugging approach that keeps the code reliable and maintainable. could you share the current codebase and the first kernel you would like optimized so i can estimate the effort and begin immediately?
$9 USD in 40 days
2.9
2.9

Hello, I am a Senior Software Engineer with experience in C++, Python, GPU computing, and AI performance optimization. I have worked with deep learning pipelines, PyTorch, CUDA-based workflows, and performance-critical systems. I can help with: CUDA kernel debugging and optimization GPU memory and execution performance tuning PyTorch CUDA extension development Tensor operations, indexing, and correctness validation Benchmarking and profiling GPU workloads Linux-based development, Git workflows, and debugging I understand the importance of first achieving correctness, then optimizing kernels for speed and efficiency. I am comfortable working through remote access environments and collaborating long-term on iterative GPU optimization tasks. I can provide reliable support for CUDA/C++ development and performance improvements.
$20 USD in 40 days
2.2
2.2

Hello, for the CUDA kernel optimization, are there specific performance metrics or targets you are aiming for? My approach involves profiling issues, then applying targeted optimizations like those that improved page load on TryReplify from 3.4s to 2.2s. I have significant experience with debugging and optimizing CUDA correctness issues. We can discuss your current needs.
$7 USD in 40 days
0.0
0.0

Hi, I see you're looking for an experienced CUDA/C++ developer for long-term remote collaboration on GPU optimization tasks. With 7+ years of experience in CUDA kernel development and performance optimization, I’m confident in my ability to help you debug and enhance your CUDA implementations effectively. My approach involves first identifying any correctness issues, then optimizing the GPU performance to ensure your project runs smoothly. I have worked on projects involving CUDA streams and memory optimization, which would be beneficial for the tasks you've outlined, such as debugging tensor indexing and convolution operations. I'm familiar with PyTorch and have experience using Linux and Git, so I can jump right into your development environment. What specific challenges have you faced with CUDA kernel correctness in your current implementations?
$8 USD in 7 days
0.0
0.0

I am an experienced CUDA and C++ developer ready to support your long-term remote project with AnyDesk or remote access. I understand the need for efficient parallel computing and robust code maintenance. I have successfully delivered high-performance CUDA applications optimizing complex algorithms and ensuring stable deployments. My expertise includes memory management and kernel optimization. I focus on clear communication and iterative progress to align with your goals. Happy to review your current setup and get this back to a stable state.
$7 USD in 7 days
0.0
0.0

Hello, I am a C++ developer with a strong foundation in modern C++, debugging, performance-oriented programming, and software development practices. I have experience building C++ applications, working with object-oriented design, memory management, data structures, and solving complex programming issues. While my main experience has been in C++ and embedded systems rather than production CUDA development, I have a strong understanding of low-level programming concepts that are directly relevant to GPU programming, including memory management, optimization, debugging, and efficient software design. My experience includes: * Modern C++ (C++17/20) * Object-Oriented Programming and clean architecture * Debugging and analyzing complex code issues * Linux development environment * Git/GitHub workflows * Performance optimization concepts * Embedded C/C++ development with STM32 and microcontrollers I am currently expanding my knowledge in GPU computing and CUDA, and I am comfortable learning and working with new low-level technologies. I approach problems carefully by understanding the root cause, reproducing issues, testing solutions, and improving performance without sacrificing correctness. I would be happy to discuss your current CUDA tasks, understand your workflow, and contribute to debugging and optimization efforts. Thank you for considering my application.
$9 USD in 40 days
0.0
0.0

Athi River, Kenya
Payment method verified
Member since Jul 7, 2024
$7-11 USD / hour
$7-10 USD / hour
$30-250 USD
₹600-1500 INR
$7-8 USD / hour
₹1250-2500 INR / hour
$2-8 USD / hour
₹1500-12500 INR
$750-1500 USD
₹750-1250 INR / hour
$250-750 USD
₹12500-37500 INR
$10-6000 CAD
$250-750 AUD
$250-750 USD
₹1500-12500 INR
$7-11 USD / hour
£75-76 GBP / hour
$8-15 CAD / hour
₹600-1500 INR
$30-250 USD
$2-8 USD / hour