AI & MLResearchers introduced MEDEM to design multi-engine deep learning chips
They claim the design methodology yields up to 4.84x better energy-delay product and 1.59x higher throughput across 51 deep learning workloads.
Papers
Preprints and accepted papers worth a scan in plain-language titles with one sentence of context, and a straight link to the source. The images are a little weird, on purpose, for funsies. (Because science is supposed to be fun). The AI bot date parser is so bad at dates, so just rejoice that at least the papers it finds are really neat ones even if they are older.
1–24 of 302 / 302 papers
AI & MLResearchers introduced MEDEM to design multi-engine deep learning chips
They claim the design methodology yields up to 4.84x better energy-delay product and 1.59x higher throughput across 51 deep learning workloads.
AI & MLResearchers scaled analog in-memory AI training to 123M parameters
They used mixed-precision analog in-memory computing to train a 123M-parameter Transformer with loss scaling comparable to digital hardware.
SimulationHPC researchers publish roadmap for mixed-precision computing
A 40-author group laid out a framework to push low-precision math in scientific computing while maintaining accuracy and cutting energy.
NetworksResearchers surveyed distributed asynchronous many-task models for HPC
A new survey analyzes runtimes like Charm++, HPX, and Legion to show where task-based models beat traditional MPI+X approaches.
AcceleratorsResearchers containerized EDA tools to boost chip simulation throughput
They report a 35% boost in simulation throughput on IBM Spectrum LSF clusters using containerized workflows and automated CI/CD pipelines.
Math & SolversResearchers reviewed floating-point errors in large-scale scientific computing
A review evaluates error mitigation methods and argues mixed-precision schedules beat all-or-nothing precision choices in scientific workloads.
Cloud & SystemsResearcher combined Prometheus, Grafana, and ELK to monitor hybrid HPC setups
The framework uses OpenTelemetry to unite metrics and logs across HPC and Kubernetes clusters, running across 150 pods with minimal overhead.
AcceleratorsResearchers open-sourced a lossless lookup table compression scheme for FPGAs
CompressedLUT uses decomposition and self-similarity to compress FPGA lookup tables for neural networks and math functions without losing accuracy.
QuantumResearchers built a gateway to route hybrid quantum workloads
They say the gateway cuts end-to-end latency by 13% to 25% on NISQ jobs compared to direct SDK calls, with sub-30 ms overhead.
Math & SolversResearchers built TC-SparIG to speed up sparse convolution on GPUs
They claim their sparse format and dataflow optimizations make sparse convolutions up to 1.63x faster than TorchSparse++ on GPUs.
AcceleratorsResearchers benchmarked OpenACC multi-GPU transfers against CUDA
A new study evaluates low-level OpenACC multi-GPU data transfer APIs against CUDA using parallel 3D sweeping algorithms.
AcceleratorsResearchers developed SmartBatchLLM to speed up LLM serving
They claim the adaptive batching scheduler reduces time to first token in vLLM by up to 58% under high concurrency without modifying CUDA kernels.
SimulationResearchers speed up planar Quickhull using SIMD and multicore tuning
Authors say VQhull delivers up to 11x parallel speedups over prior art while hitting up to 100% of peak memory bandwidth on non-NUMA CPUs.
Data & I/OEngineers automated pre-silicon testing for DDR5 memory controllers
They built a framework that pulls timing data directly from verification files to test memory controller RTL without recompilation.
Math & SolversResearchers designed GEM-KMeans for memory-efficient clustering on GPUs
By fusing updates into matrix multiplication epilogues, the method cuts memory use down to a single factor array in HBM.
AcceleratorsResearchers built Wavel to compile code for wafer-scale chips
They claim it boosts throughput up to 1.35x over WaferLLM on Cerebras WSE-3 hardware by optimizing physical placement and execution schedules.
AcceleratorsResearchers built MorphX to dynamically resize GPU kernels
They claim MorphX boosts background task throughput by 2.24x without hurting foreground latency by letting kernels yield GPU compute on the fly.
AcceleratorsResearchers built StreamInfer to cut MoE communication barriers
They claim barrier-free token streaming boosts MoE decoding throughput up to 1.7x on 16-GPU A100 and L40S clusters.
AcceleratorsAnchor speeds up GPU crash recovery by decoupling memory ownership
They claim decoupling GPU memory via a daemon cuts recovery time by 60.7% for inference and 26.6% for training after shallow process crashes.
AcceleratorsResearchers built Meld to fuse GPU kernels across SMs
They claim the framework speeds up dynamic workloads like decoding attention by an average of 1.3x.
SimulationResearchers built DDB to enable source-level debugging across distributed apps
They say DDB virtualizes process clocks to avoid timeout cascades while adding only 1 to 5 percent throughput overhead across 122 processes.
SimulationResearchers verify deterministic parallel execution with DeterV
The runtime uses Verus to verify equivalence to serial execution, achieving a 3:1 proof-to-code ratio with minimal runtime overhead.
NetworksResearchers propose hybrid load balancer for virtualized HPC
The method mixes mesh migration with temporary precision reduction, hitting a 1.215x peak speedup at eight MPI ranks on a virtualized testbed.
AI & MLResearchers benchmarked CXL memory expansion for AI inference on Xeon
They found adding CXL Type-3 memory boosted synthetic bandwidth by up to 40% and LLaMA-13B inference throughput by up to 20%.