Vendor Updates
Vendor Updates
Everything the vendors have put out. Pick a vendor to catch up on their run of announcements.
Sep 29, 20261
- SK hynix · 1d agoSK hynix displayed HBM4 memory for Nvidia Rubin
SK hynix showcased 12-layer 36GB HBM4 and SOCAMM2 memory modules built for Nvidia's Vera Rubin superchip.
Read originalnews.skhynix.com/en/tsmc-oip-conference-2026SK hynix · 1d ago
SK hynix displayed HBM4 memory for Nvidia Rubin
SK hynix showcased 12-layer 36GB HBM4 and SOCAMM2 memory modules built for Nvidia's Vera Rubin superchip.
Read original https://news.skhynix.com/en/tsmc-oip-conference-2026hbmmemorypackaginghbmmemorypackaging
Sep 28, 20263
- Cerebras · 2d agoGimlet Labs adds Cerebras wafer-scale chips to its inference cloud
They claim combining Cerebras wafer-scale chips with GPUs will deliver up to 3,000 tokens per second for AI inference workloads.
Read originalcerebras.ai/press-release/gimlet-labs-adds-cerebras-to-deliver-ultrafast-ai-inference-through-gimlet-cloud-deploymentCerebras · 2d ago
Gimlet Labs adds Cerebras wafer-scale chips to its inference cloud
They claim combining Cerebras wafer-scale chips with GPUs will deliver up to 3,000 tokens per second for AI inference workloads.
Read original https://cerebras.ai/press-release/gimlet-labs-adds-cerebras-to-deliver-ultrafast-ai-inference-through-gimlet-cloud-deploymentno tagsno tags - Crusoe · 2d agoCrusoe is building an AI data center campus for Google in Texas
The facility is co-located with 531 MW of wind generation and uses a closed-loop, non-evaporative liquid cooling system.
Read originalcrusoe.ai/resources/newsroom/crusoe-developing-data-center-campus-in-armstrong-county-texasCrusoe · 2d ago
Crusoe is building an AI data center campus for Google in Texas
The facility is co-located with 531 MW of wind generation and uses a closed-loop, non-evaporative liquid cooling system.
Read original https://crusoe.ai/resources/newsroom/crusoe-developing-data-center-campus-in-armstrong-county-texasdatacenterliquid-coolingclouddatacenterliquid-coolingcloud - AMD · 2d agoAMD agrees to buy AI startup World Labs for $8.2 billion
AMD agreed to buy AI startup World Labs in an $8.2 billion all-stock deal to help shape its future AI hardware and software roadmaps.
Read originalir.amd.com/news-events/press-releases/detail/1299/amd-to-acquire-world-labs-to-advance-the-future-of-ai-computeAMD · 2d ago
AMD agrees to buy AI startup World Labs for $8.2 billion
AMD agreed to buy AI startup World Labs in an $8.2 billion all-stock deal to help shape its future AI hardware and software roadmaps.
Read original https://ir.amd.com/news-events/press-releases/detail/1299/amd-to-acquire-world-labs-to-advance-the-future-of-ai-computeaiacquisitioncomputeaiacquisitioncompute
Sep 27, 20261
- Lambda · 3d agoLambda plans new AI data center in Oklahoma
The company will build an AI cloud facility in Mayes County with closed-loop cooling and expects to generate $500 million in local taxes over ten years.
Read originallambda.ai/blog/new-lambda-data-center-in-oklahomaLambda · 3d ago
Lambda plans new AI data center in Oklahoma
The company will build an AI cloud facility in Mayes County with closed-loop cooling and expects to generate $500 million in local taxes over ten years.
Read original https://lambda.ai/blog/new-lambda-data-center-in-oklahomaclouddata-centergpuclouddata-centergpu
Sep 26, 20261
- hgpu.org · 4d agohgpu.org updated its GPU computing research paper directory
The repository indexed recent work on Triton kernel optimization, NVIDIA Blackwell confidential computing, and multi-GPU parallelism.
Read originalhgpu.org/?cat=11hgpu.org · 4d ago
hgpu.org updated its GPU computing research paper directory
The repository indexed recent work on Triton kernel optimization, NVIDIA Blackwell confidential computing, and multi-GPU parallelism.
Read original https://hgpu.org/?cat=11no tagsno tags
Sep 23, 20264
- Ionq · Sep 23, 2026IonQ will install its Superion 256 QPU at Nvidia's quantum research center
The 256-qubit system will link directly to an Nvidia GB200 NVL72 cluster over NVQLink next year.
Read originalionq.com/news/ionq-to-advance-quantum-supercomputing-by-bringing-first-qpu-to-nvidia-accelerated-quantum-research-centerIonq · Sep 23, 2026
IonQ will install its Superion 256 QPU at Nvidia's quantum research center
The 256-qubit system will link directly to an Nvidia GB200 NVL72 cluster over NVQLink next year.
Read original https://ionq.com/news/ionq-to-advance-quantum-supercomputing-by-bringing-first-qpu-to-nvidia-accelerated-quantum-research-centerno tagsno tags - VAST Data · Sep 23, 2026CINECA retools supercomputer storage to handle AI alongside traditional HPC
CINECA says AI now makes up over half of Leonardo's workload, driving a shift to multi-protocol data platforms across its clusters.
Read originalvastdata.com/blog/cineca-rebuilding-hpc-for-aiVAST Data · Sep 23, 2026
CINECA retools supercomputer storage to handle AI alongside traditional HPC
CINECA says AI now makes up over half of Leonardo's workload, driving a shift to multi-protocol data platforms across its clusters.
Read original https://vastdata.com/blog/cineca-rebuilding-hpc-for-aistoragehpcsupercomputerstoragehpcsupercomputer - Supermicro · Sep 23, 2026Supermicro started shipping liquid-cooled Nvidia Vera Rubin NVL72 racks
Each 72-GPU rack provides 20.7 TB of HBM4 and 216 TB/s of NVLink bandwidth across 18 compute trays with direct liquid cooling.
Read originalsupermicro.com/en/pressreleases/supermicro-now-shipping-nvidia-vera-rubin-nvl72-racksSupermicro · Sep 23, 2026
Supermicro started shipping liquid-cooled Nvidia Vera Rubin NVL72 racks
Each 72-GPU rack provides 20.7 TB of HBM4 and 216 TB/s of NVLink bandwidth across 18 compute trays with direct liquid cooling.
Read original https://supermicro.com/en/pressreleases/supermicro-now-shipping-nvidia-vera-rubin-nvl72-racksgpuserveraidatacentergpuserveraidatacenter - CoreWeave · Sep 23, 2026CoreWeave earned a third Platinum cluster rating from SemiAnalysis
SemiAnalysis gave CoreWeave its top cluster rating again, citing custom tools for straggler detection, node validation, and GPU testing.
Read originalcoreweave.com/news/coreweave-becomes-the-only-provider-to-earn-semianalysis-platinum-clustermax-tm-rating-three-consecutive-timesCoreWeave · Sep 23, 2026
CoreWeave earned a third Platinum cluster rating from SemiAnalysis
SemiAnalysis gave CoreWeave its top cluster rating again, citing custom tools for straggler detection, node validation, and GPU testing.
Read original https://coreweave.com/news/coreweave-becomes-the-only-provider-to-earn-semianalysis-platinum-clustermax-tm-rating-three-consecutive-timesbenchmarkclustercloud-computingbenchmarkclustercloud-computing
Sep 22, 20266
- CoreWeave · Sep 22, 2026CoreWeave added cross-region writes to its AI object storage
They added local-latency cross-region writes to reduce checkpoint pauses when training on distributed GPU clusters.
Read originalcoreweave.com/blog/new-in-coreweave-ai-object-storage-cross-region-writes-and-archive-storageCoreWeave · Sep 22, 2026
CoreWeave added cross-region writes to its AI object storage
They added local-latency cross-region writes to reduce checkpoint pauses when training on distributed GPU clusters.
Read original https://coreweave.com/blog/new-in-coreweave-ai-object-storage-cross-region-writes-and-archive-storageno tagsno tags - Lgcorp · Sep 22, 2026LG qualified a 2.6-megawatt coolant distribution unit for Nvidia AI clusters
LG says its 2.6 MW liquid-cooling unit met Nvidia DSX reference design requirements for dense AI server clusters.
Read originallgcorp.com/media/release/30597Lgcorp · Sep 22, 2026
LG qualified a 2.6-megawatt coolant distribution unit for Nvidia AI clusters
LG says its 2.6 MW liquid-cooling unit met Nvidia DSX reference design requirements for dense AI server clusters.
Read original https://lgcorp.com/media/release/30597no tagsno tags - VAST Data · Sep 22, 2026VAST Data launched DataEnclave for confidential AI workloads
VAST says DataEnclave runs AI workloads across encrypted CPU and GPU memory without exposing weights or data to infrastructure operators.
Read originalvastdata.com/blog/introducing-vast-dataenclave-confidential-ai-for-sensitive-data-and-proprietary-modelsVAST Data · Sep 22, 2026
VAST Data launched DataEnclave for confidential AI workloads
VAST says DataEnclave runs AI workloads across encrypted CPU and GPU memory without exposing weights or data to infrastructure operators.
Read original https://vastdata.com/blog/introducing-vast-dataenclave-confidential-ai-for-sensitive-data-and-proprietary-modelsstorageaisecurityconfidential-computingstorageaisecurityconfidential-computing - CoreWeave · Sep 22, 2026CoreWeave posted MLPerf Inference 6.1 results on Nvidia Blackwell Ultra
They claim a single GB300 NVL72 rack reached over 1.16 million tokens per second running GPT-OSS-120B in server mode.
Read originalcoreweave.com/blog/coreweave-leads-cloud-providers-in-mlperf-r-inference-v6-1-performance-with-nvidia-blackwell-ultraCoreWeave · Sep 22, 2026
CoreWeave posted MLPerf Inference 6.1 results on Nvidia Blackwell Ultra
They claim a single GB300 NVL72 rack reached over 1.16 million tokens per second running GPT-OSS-120B in server mode.
Read original https://coreweave.com/blog/coreweave-leads-cloud-providers-in-mlperf-r-inference-v6-1-performance-with-nvidia-blackwell-ultrabenchmarkmlperfgpucloudbenchmarkmlperfgpucloud - SK hynix · Sep 22, 2026SK hynix discussed AI memory roadmaps at its global forum
SK hynix executives discussed their HBM roadmap and expansion of production hubs in Korea and the US to meet AI memory demand.
Read originalnews.skhynix.com/en/2026-global-forumSK hynix · Sep 22, 2026
SK hynix discussed AI memory roadmaps at its global forum
SK hynix executives discussed their HBM roadmap and expansion of production hubs in Korea and the US to meet AI memory demand.
Read original https://news.skhynix.com/en/2026-global-forumno tagsno tags - CoreWeave · Sep 22, 2026CoreWeave detailed bringing up multi-rack Nvidia Rubin NVL72 clusters
CoreWeave detailed its automated lifecycle tools for managing liquid cooling, power, and firmware across multi-rack Nvidia Vera Rubin NVL72 clusters.
Read originalcoreweave.com/blog/what-it-takes-to-bring-up-a-multi-rack-nvidia-vera-rubin-nvl72-clusterCoreWeave · Sep 22, 2026
CoreWeave detailed bringing up multi-rack Nvidia Rubin NVL72 clusters
CoreWeave detailed its automated lifecycle tools for managing liquid cooling, power, and firmware across multi-rack Nvidia Vera Rubin NVL72 clusters.
Read original https://coreweave.com/blog/what-it-takes-to-bring-up-a-multi-rack-nvidia-vera-rubin-nvl72-clustergpuclusterdatacenterinterconnectgpuclusterdatacenterinterconnect
Sep 20, 20263
- Supercomputing News · Sep 20, 2026Rambus released a security subsystem to pair with Caliptra chips
Rambus says its new hardware security block runs alongside an unmodified open-source Caliptra core to preserve trademark eligibility.
Read originalsupercomputing.news/ai/rambus-caliptra-root-of-trust-beside-open-coreSupercomputing News · Sep 20, 2026
Rambus released a security subsystem to pair with Caliptra chips
Rambus says its new hardware security block runs alongside an unmodified open-source Caliptra core to preserve trademark eligibility.
Read original https://supercomputing.news/ai/rambus-caliptra-root-of-trust-beside-open-coresecurityhardwaredatacentersecurityhardwaredatacenter - hgpu.org · Sep 20, 2026Researchers benchmarked LLM prefix reuse on NVIDIA H100 GPUs
PrefixBench-H100 evaluates vLLM and TensorRT-LLM on H100 GPUs to map where KV-cache prefix reuse lowers first-token latency and where cache pressure erodes it.
Read originalhgpu.org/?p=31227hgpu.org · Sep 20, 2026
Researchers benchmarked LLM prefix reuse on NVIDIA H100 GPUs
PrefixBench-H100 evaluates vLLM and TensorRT-LLM on H100 GPUs to map where KV-cache prefix reuse lowers first-token latency and where cache pressure erodes it.
Read original https://hgpu.org/?p=31227gpubenchmarkllmgpubenchmarkllm - Samsung Semiconductor · Sep 20, 2026Samsung introduced EquFlashV2 to speed up atomistic simulations
They claim the GPU-optimized ML force field scales at O(N) instead of O(N³) to replace costly DFT in semiconductor material design.
Read originalsemiconductor.samsung.com/kr/news-events/tech-blog/beyond-dfts-computational-limits-equflashv2-and-the-future-of-ai-driven-semiconductor-materials-simulationSamsung Semiconductor · Sep 20, 2026
Samsung introduced EquFlashV2 to speed up atomistic simulations
They claim the GPU-optimized ML force field scales at O(N) instead of O(N³) to replace costly DFT in semiconductor material design.
Read original https://semiconductor.samsung.com/kr/news-events/tech-blog/beyond-dfts-computational-limits-equflashv2-and-the-future-of-ai-driven-semiconductor-materials-simulationmaterials-sciencesimulationmachine-learningmaterials-sciencesimulationmachine-learning
Sep 18, 20261
- SK hynix · Sep 18, 2026SK hynix showed High Bandwidth Flash and PIM cards for AI inference
They presented High Bandwidth Flash NAND and AiMX processing-in-memory accelerator cards designed to reduce memory bottlenecks in AI inference.
Read originalnews.skhynix.com/en/ai-infra-summit-2026SK hynix · Sep 18, 2026
SK hynix showed High Bandwidth Flash and PIM cards for AI inference
They presented High Bandwidth Flash NAND and AiMX processing-in-memory accelerator cards designed to reduce memory bottlenecks in AI inference.
Read original https://news.skhynix.com/en/ai-infra-summit-2026no tagsno tags
Sep 17, 20263
- Marvell · Sep 17, 2026GlobalFoundries expanded SiGe manufacturing deal with Marvell for optics
The multi-year deal increases Vermont fab capacity for 200G-per-lane SiGe chips used in optical transceivers and co-packaged optics.
Read originalgf.com/news-and-events/news/globalfoundries-and-marvell-expand-collaboration-fornext-generation-optical-connectivityMarvell · Sep 17, 2026
GlobalFoundries expanded SiGe manufacturing deal with Marvell for optics
The multi-year deal increases Vermont fab capacity for 200G-per-lane SiGe chips used in optical transceivers and co-packaged optics.
Read original https://gf.com/news-and-events/news/globalfoundries-and-marvell-expand-collaboration-fornext-generation-optical-connectivityno tagsno tags - Crusoe · Sep 17, 2026Crusoe raised $3.9B to expand AI datacenter infrastructure
The funding round values Crusoe at $30.9B as they expand power generation, modular data center builds, and cloud capacity for AI workloads.
Read originalcrusoe.ai/resources/newsroom/crusoe-announces-series-f-fundingCrusoe · Sep 17, 2026
Crusoe raised $3.9B to expand AI datacenter infrastructure
The funding round values Crusoe at $30.9B as they expand power generation, modular data center builds, and cloud capacity for AI workloads.
Read original https://crusoe.ai/resources/newsroom/crusoe-announces-series-f-fundingdatacentercloudaidatacentercloudai - Amazon · Sep 17, 2026AWS added bulk job cancellation and termination to AWS Batch
Users can now cancel or terminate up to 50 individual or array jobs in a single API call across all AWS Batch regions.
Read originalaws.amazon.com/about-aws/whats-new/2026/09/aws-batch-bulk-cancellationAmazon · Sep 17, 2026
AWS added bulk job cancellation and termination to AWS Batch
Users can now cancel or terminate up to 50 individual or array jobs in a single API call across all AWS Batch regions.
Read original https://aws.amazon.com/about-aws/whats-new/2026/09/aws-batch-bulk-cancellationno tagsno tags
Sep 16, 20262
- nvi · Sep 16, 2026Nvidia, Google, and Emerald AI formed the AI Energy Management Alliance
The group aims to standardize how AI data centers adjust electricity demand in real time to speed up power grid interconnections.
Read originalblogs.nvidia.com/blog/ai-energy-management-alliancenvi · Sep 16, 2026
Nvidia, Google, and Emerald AI formed the AI Energy Management Alliance
The group aims to standardize how AI data centers adjust electricity demand in real time to speed up power grid interconnections.
Read original https://blogs.nvidia.com/blog/ai-energy-management-allianceno tagsno tags - Samsung Semiconductor · Sep 16, 2026Samsung detailed a 245TB PCIe 5.0 QLC SSD for AI data centers
Samsung says the E3.S drive hits 245.76TB with 4.5 GB/s sequential writes and a 0.6 drive writes per day endurance rating.
Read originalsemiconductor.samsung.com/news-events/tech-blog/samsung-bm1773-ultra-high-density-storage-solution-for-ai-data-centersSamsung Semiconductor · Sep 16, 2026
Samsung detailed a 245TB PCIe 5.0 QLC SSD for AI data centers
Samsung says the E3.S drive hits 245.76TB with 4.5 GB/s sequential writes and a 0.6 drive writes per day endurance rating.
Read original https://semiconductor.samsung.com/news-events/tech-blog/samsung-bm1773-ultra-high-density-storage-solution-for-ai-data-centersssdstorageflash-memoryai-infrastructuressdstorageflash-memoryai-infrastructure
Sep 15, 20264
- Astera Labs · Sep 15, 2026Astera Labs introduced memory controllers for AI fabrics and CXL
They claim the controllers double bandwidth over prior parts and improve inference time to first token by up to 62% using fabric-attached memory.
Read originalasteralabs.gcs-web.com/news-releases/news-release-details/astera-labs-bolsters-leo-smart-memory-controller-family-agenticAstera Labs · Sep 15, 2026
Astera Labs introduced memory controllers for AI fabrics and CXL
They claim the controllers double bandwidth over prior parts and improve inference time to first token by up to 62% using fabric-attached memory.
Read original https://asteralabs.gcs-web.com/news-releases/news-release-details/astera-labs-bolsters-leo-smart-memory-controller-family-agenticno tagsno tags - SiPearl · Sep 15, 2026SiFive and AMD demonstrated ROCm on RISC-V datacenter servers
They ran a Gemma model using ROCm 10 on a 32-core P870-D host offloading inference to AMD Radeon AI PRO R9700 GPUs.
Read originalsifive.com/press/sifive-amd-rocm-riscv-datacenter-serversSiPearl · Sep 15, 2026
SiFive and AMD demonstrated ROCm on RISC-V datacenter servers
They ran a Gemma model using ROCm 10 on a 32-core P870-D host offloading inference to AMD Radeon AI PRO R9700 GPUs.
Read original https://sifive.com/press/sifive-amd-rocm-riscv-datacenter-serversno tagsno tags - Micron · Sep 15, 2026Micron demoed 512GB DDR5 RDIMMs reaching 9,200 MT/s
Micron says the TSV-stacked modules enable up to 12TB per dual-socket server and cut operating power by over 60% compared to four 128GB DIMMs.
Read originalinvestors.micron.com/news/press-release/2026/Micron-Advances-Memory-Innovation-With-the-Worlds-First-Ultra-Dense-Module-for-Next-Generation-Servers/default.aspxMicron · Sep 15, 2026
Micron demoed 512GB DDR5 RDIMMs reaching 9,200 MT/s
Micron says the TSV-stacked modules enable up to 12TB per dual-socket server and cut operating power by over 60% compared to four 128GB DIMMs.
Read original https://investors.micron.com/news/press-release/2026/Micron-Advances-Memory-Innovation-With-the-Worlds-First-Ultra-Dense-Module-for-Next-Generation-Servers/default.aspxdrammemorydatacenterserverdrammemorydatacenterserver - Astera Labs · Sep 15, 2026Astera Labs launched Leo memory controllers for AI fabrics and CXL pooling
They claim the controllers double memory bandwidth and cut time to first token by up to 62% in AI inference workloads.
Read originalasteralabs.com/news/astera-labs-bolsters-leo-smart-memory-controller-family-for-agentic-ai-and-general-purpose-cloud-workloadsAstera Labs · Sep 15, 2026
Astera Labs launched Leo memory controllers for AI fabrics and CXL pooling
They claim the controllers double memory bandwidth and cut time to first token by up to 62% in AI inference workloads.
Read original https://asteralabs.com/news/astera-labs-bolsters-leo-smart-memory-controller-family-for-agentic-ai-and-general-purpose-cloud-workloadscxlmemoryinterconnectcxlmemoryinterconnect
Sep 14, 20263
- NVIDIA · Sep 14, 2026Nvidia updated CUDA-Q for fault-tolerant quantum computing
Nvidia added CUDA-Q Logical to its hybrid quantum platform, which Fermilab claims sped up fault-tolerant architecture exploration by 7x.
Read originalnvidianews.nvidia.com/news/nvidia-expands-open-source-cuda-q-platform-for-fault-tolerant-quantum-computingNVIDIA · Sep 14, 2026
Nvidia updated CUDA-Q for fault-tolerant quantum computing
Nvidia added CUDA-Q Logical to its hybrid quantum platform, which Fermilab claims sped up fault-tolerant architecture exploration by 7x.
Read original https://nvidianews.nvidia.com/news/nvidia-expands-open-source-cuda-q-platform-for-fault-tolerant-quantum-computingquantumcudahpcquantumcudahpc - Lambda · Sep 14, 2026Lambda signed the White House Ratepayer Protection Pledge
The AI cloud provider agreed to principles requiring data center builders to cover their own power use and protect local electric grids.
Read originallambda.ai/blog/lambda-signs-the-white-house-ratepayer-protection-pledgeLambda · Sep 14, 2026
Lambda signed the White House Ratepayer Protection Pledge
The AI cloud provider agreed to principles requiring data center builders to cover their own power use and protect local electric grids.
Read original https://lambda.ai/blog/lambda-signs-the-white-house-ratepayer-protection-pledgeno tagsno tags - DDN · Sep 14, 2026DDN won an industry storage award for EXAScaler
DDN says its parallel storage platform currently supports over one million GPUs in AI and supercomputing environments.
Read originalddn.com/press-releases/ddn-exascaler-wins-2026-siliconangle-techforward-award-for-data-storage-managementDDN · Sep 14, 2026
DDN won an industry storage award for EXAScaler
DDN says its parallel storage platform currently supports over one million GPUs in AI and supercomputing environments.
Read original https://ddn.com/press-releases/ddn-exascaler-wins-2026-siliconangle-techforward-award-for-data-storage-managementno tagsno tags
Sep 13, 20263
- Apple · Sep 13, 2026Researchers compared compute performance of diffusion and autoregressive LLMs
They found diffusion models boost arithmetic intensity via parallel tokens but lag autoregressive models on long contexts and batched throughput.
Read originalmachinelearning.apple.com/research/diffusion-autoregressive-performanceApple · Sep 13, 2026
Researchers compared compute performance of diffusion and autoregressive LLMs
They found diffusion models boost arithmetic intensity via parallel tokens but lag autoregressive models on long contexts and batched throughput.
Read original https://machinelearning.apple.com/research/diffusion-autoregressive-performancebenchmarkdiffusion-modelllmbenchmarkdiffusion-modelllm - NVIDIA · Sep 13, 2026Nvidia published a validator tool for NVCF GPU clusters
The container tests control plane health, GPU availability, network policies, and storage drivers on clusters running NVCF workloads.
Read originalcatalog.ngc.nvidia.com/orgs/nvidia/nvcf-byoc/containers/cluster-validatorNVIDIA · Sep 13, 2026
Nvidia published a validator tool for NVCF GPU clusters
The container tests control plane health, GPU availability, network policies, and storage drivers on clusters running NVCF workloads.
Read original https://catalog.ngc.nvidia.com/orgs/nvidia/nvcf-byoc/containers/cluster-validatorno tagsno tags - Apple · Sep 13, 2026Apple and Berkeley team introduced Arbitrage for faster LLM reasoning
They claim the step-level speculative decoding framework cuts LLM reasoning inference latency by up to 2x at matched accuracy.
Read originalmachinelearning.apple.com/research/arbitrage-efficient-reasoningApple · Sep 13, 2026
Apple and Berkeley team introduced Arbitrage for faster LLM reasoning
They claim the step-level speculative decoding framework cuts LLM reasoning inference latency by up to 2x at matched accuracy.
Read original https://machinelearning.apple.com/research/arbitrage-efficient-reasoningmachine-learninginferencellmmachine-learninginferencellm
Sep 12, 20262
- Apple · Sep 12, 2026Apple tested an enlarge-and-prune pipeline for LLM pretraining
Researchers say pretraining an enlarged 2.8B model and pruning it to 1.3B across 2T tokens beats training the target size from scratch.
Read originalmachinelearning.apple.com/research/idea-prune-pipelineApple · Sep 12, 2026
Apple tested an enlarge-and-prune pipeline for LLM pretraining
Researchers say pretraining an enlarged 2.8B model and pruning it to 1.3B across 2T tokens beats training the target size from scratch.
Read original https://machinelearning.apple.com/research/idea-prune-pipelinellmmodel-pruningdeep-learningllmmodel-pruningdeep-learning - Apple · Sep 12, 2026Apple modeled data mixture scaling laws across 2,000 LLM training runs
They found scarce target corpora can be reused 15 to 20 times during mixture pretraining before hitting diminishing returns.
Read originalmachinelearning.apple.com/research/scaling-laws-mixture-pretrainingApple · Sep 12, 2026
Apple modeled data mixture scaling laws across 2,000 LLM training runs
They found scarce target corpora can be reused 15 to 20 times during mixture pretraining before hitting diminishing returns.
Read original https://machinelearning.apple.com/research/scaling-laws-mixture-pretrainingscaling-lawllmpretrainingscaling-lawllmpretraining
Sep 11, 20261
- Arm · Sep 11, 2026Arm detailed AI infrastructure and developer tools at Arm Create
At events in China, Arm highlighted cloud options like Neoverse CSS N4 alongside profiling tools to help engineers deploy AI models across edge and cloud.
Read originalnewsroom.arm.com/blog/takeaways-from-arm-create-china-2026Arm · Sep 11, 2026
Arm detailed AI infrastructure and developer tools at Arm Create
At events in China, Arm highlighted cloud options like Neoverse CSS N4 alongside profiling tools to help engineers deploy AI models across edge and cloud.
Read original https://newsroom.arm.com/blog/takeaways-from-arm-create-china-2026no tagsno tags
Sep 10, 20262
- SK hynix · Sep 10, 2026SK hynix discussed AI memory strategy at its annual tech forum
SK hynix executives met with AI leaders to discuss shifting workload demands and future directions for memory development.
Read originalnews.skhynix.com/en/future-forum-2026SK hynix · Sep 10, 2026
SK hynix discussed AI memory strategy at its annual tech forum
SK hynix executives met with AI leaders to discuss shifting workload demands and future directions for memory development.
Read original https://news.skhynix.com/en/future-forum-2026no tagsno tags - Groq · Sep 10, 2026Groq added a batch processing API to GroqCloud
They offer 24-hour turnaround on bulk inference jobs for models like Llama 3.3 70B at a 25% discount off on-demand rates.
Read originalgroq.com/blog/batch-processing-with-groqcloud-for-ai-inference-workloadsGroq · Sep 10, 2026
Groq added a batch processing API to GroqCloud
They offer 24-hour turnaround on bulk inference jobs for models like Llama 3.3 70B at a 25% discount off on-demand rates.
Read original https://groq.com/blog/batch-processing-with-groqcloud-for-ai-inference-workloadscloudinferencebatch-processingcloudinferencebatch-processing
Sep 8, 20269
- Arm · Sep 8, 2026Arm introduced Neoverse CSS N4 and an AGI CPU for cloud AI workloads
They announced Neoverse CSS N4 for scale-out servers alongside a new AGI CPU designed for datacenter AI infrastructure.
Read originalnewsroom.arm.com/news/arm-everywhere-the-compute-platform-for-agentic-aiArm · Sep 8, 2026
Arm introduced Neoverse CSS N4 and an AGI CPU for cloud AI workloads
They announced Neoverse CSS N4 for scale-out servers alongside a new AGI CPU designed for datacenter AI infrastructure.
Read original https://newsroom.arm.com/news/arm-everywhere-the-compute-platform-for-agentic-aino tagsno tags - Samsung Semiconductor · Sep 8, 2026Samsung partnered with Mistral AI to build custom models for chip design
Samsung invested in Mistral AI to develop on-premises models tailored for semiconductor design, defect prediction, and fab optimization.
Read originalsemiconductor.samsung.com/kr/news-events/news/samsung-and-mistral-ai-announce-strategic-partnership-for-intelligence-driven-semiconductor-infrastructureSamsung Semiconductor · Sep 8, 2026
Samsung partnered with Mistral AI to build custom models for chip design
Samsung invested in Mistral AI to develop on-premises models tailored for semiconductor design, defect prediction, and fab optimization.
Read original https://semiconductor.samsung.com/kr/news-events/news/samsung-and-mistral-ai-announce-strategic-partnership-for-intelligence-driven-semiconductor-infrastructureno tagsno tags - GitHub · Sep 8, 2026NVIDIA engineers shared a tutorial on NCCL device APIs
The tutorial covers NCCL device APIs for PyTorch workloads and provides companion code exercises on GitHub.
Read originalgithub.com/jeffhammond/pytorch-china-2026-nccl-tutorialGitHub · Sep 8, 2026
NVIDIA engineers shared a tutorial on NCCL device APIs
The tutorial covers NCCL device APIs for PyTorch workloads and provides companion code exercises on GitHub.
Read original https://github.com/jeffhammond/pytorch-china-2026-nccl-tutorialnvidiancclnvidianccl - Arm · Sep 8, 2026Arm launched the Arm AI Portal for hardware-optimized models
The hub provides pre-optimized AI models for Arm silicon, claiming over 4x speedups for models like Qwen3-TTS using SME2 vector extensions.
Read originalnewsroom.arm.com/news/arm-unveils-arm-ai-portalArm · Sep 8, 2026
Arm launched the Arm AI Portal for hardware-optimized models
The hub provides pre-optimized AI models for Arm silicon, claiming over 4x speedups for models like Qwen3-TTS using SME2 vector extensions.
Read original https://newsroom.arm.com/news/arm-unveils-arm-ai-portalno tagsno tags - Qualcomm · Sep 8, 2026Qualcomm and Amazon partner on custom AI chips and 1.6T optics
The companies say they will develop custom AI inference silicon and optical interconnects running up to 1.6T for AWS data centers.
Read originalqualcomm.com/news/releases/2026/09/qualcomm-announces-multi-generational-product-collaboration-withQualcomm · Sep 8, 2026
Qualcomm and Amazon partner on custom AI chips and 1.6T optics
The companies say they will develop custom AI inference silicon and optical interconnects running up to 1.6T for AWS data centers.
Read original https://qualcomm.com/news/releases/2026/09/qualcomm-announces-multi-generational-product-collaboration-withno tagsno tags - Qoro · Sep 8, 2026Qoro and Hartree Centre partner to link quantum software with HPC
They will integrate Qoro's software with Hartree's Slurm scheduler via QRMI to build hybrid quantum-classical chemistry workflows.
Read originalqoroquantum.net/?news=qoro-signs-mou-with-hartree-centre-to-build-uk-quantum-hpc-demonstratorQoro · Sep 8, 2026
Qoro and Hartree Centre partner to link quantum software with HPC
They will integrate Qoro's software with Hartree's Slurm scheduler via QRMI to build hybrid quantum-classical chemistry workflows.
Read original https://qoroquantum.net/?news=qoro-signs-mou-with-hartree-centre-to-build-uk-quantum-hpc-demonstratorno tagsno tags - Arm · Sep 8, 2026Arm announced the Neoverse CSS N4 server subsystem
Arm claims the N3P-based subsystem scales up to 128 cores at 3.8 GHz with a 2x speed boost over the previous generation.
Read originalhardwarebrief.com/article/arm-unveils-neoverse-css-n4-server-subsystem-c4135bcdArm · Sep 8, 2026
Arm announced the Neoverse CSS N4 server subsystem
Arm claims the N3P-based subsystem scales up to 128 cores at 3.8 GHz with a 2x speed boost over the previous generation.
Read original https://hardwarebrief.com/article/arm-unveils-neoverse-css-n4-server-subsystem-c4135bcdarmservercpuprocessorarmservercpuprocessor - NVIDIA · Sep 8, 2026Nvidia introduced CUDA Rust with SIMT and Tile kernel backends
They introduced cuda-oxide for SIMT kernels via a custom codegen backend and cutile-rs for tile-based GPU programming on stable Rust.
Read originaldeveloper.nvidia.com/blog/introducing-cuda-rust-two-tracks-for-writing-gpu-kernelsNVIDIA · Sep 8, 2026
Nvidia introduced CUDA Rust with SIMT and Tile kernel backends
They introduced cuda-oxide for SIMT kernels via a custom codegen backend and cutile-rs for tile-based GPU programming on stable Rust.
Read original https://developer.nvidia.com/blog/introducing-cuda-rust-two-tracks-for-writing-gpu-kernelsno tagsno tags - Samsung Semiconductor · Sep 8, 2026Samsung invested in Mistral AI to use on-premises models in chip fabs
Samsung led Mistral's Series D round and plans to run custom models on-prem for defect detection and equipment tuning in its fabs.
Read originalsemiconductor.samsung.com/news-events/news/samsung-and-mistral-ai-announce-strategic-partnership-for-intelligence-driven-semiconductor-infrastructureSamsung Semiconductor · Sep 8, 2026
Samsung invested in Mistral AI to use on-premises models in chip fabs
Samsung led Mistral's Series D round and plans to run custom models on-prem for defect detection and equipment tuning in its fabs.
Read original https://semiconductor.samsung.com/news-events/news/samsung-and-mistral-ai-announce-strategic-partnership-for-intelligence-driven-semiconductor-infrastructureno tagsno tags
Sep 7, 20263
- Samsung · Sep 7, 2026Samsung moved half of its 4nm fab capacity to HBM4 base dies
Samsung is putting half its 4nm capacity toward HBM4 base dies to handle custom logic like on-die memory controllers and PHYs.
Read originalhardwarebrief.com/article/samsung-allocates-50-of-4nm-capacity-to-hbm4-9c36b1b6Samsung · Sep 7, 2026
Samsung moved half of its 4nm fab capacity to HBM4 base dies
Samsung is putting half its 4nm capacity toward HBM4 base dies to handle custom logic like on-die memory controllers and PHYs.
Read original https://hardwarebrief.com/article/samsung-allocates-50-of-4nm-capacity-to-hbm4-9c36b1b6hbmmemoryfoundrysemiconductorhbmmemoryfoundrysemiconductor - Intel · Sep 7, 2026Intel and ASML reported progress on High-NA EUV manufacturing
Intel says it has processed over one million wafers with High-NA EUV tools across R&D and production layers for Panther Lake chips.
Read originalintel.com/content/www/us/en/newsroom/news/intel-foundry/intel-foundry-asml-accelerate-industry-readiness-for-high-na-euv.htmlIntel · Sep 7, 2026
Intel and ASML reported progress on High-NA EUV manufacturing
Intel says it has processed over one million wafers with High-NA EUV tools across R&D and production layers for Panther Lake chips.
Read original https://intel.com/content/www/us/en/newsroom/news/intel-foundry/intel-foundry-asml-accelerate-industry-readiness-for-high-na-euv.htmlno tagsno tags - Samsung Semiconductor · Sep 7, 2026Samsung plans to use ASML High NA EUV for DRAM by 2028
Samsung says it will adopt ASML's High NA EUV lithography for mass DRAM production by 2028 and back a move to 12-inch photomasks.
Read originalsemiconductor.samsung.com/news-events/news/samsung-electronics-and-asml-expand-strategic-collaboration-for-next-generation-semiconductor-manufacturingSamsung Semiconductor · Sep 7, 2026
Samsung plans to use ASML High NA EUV for DRAM by 2028
Samsung says it will adopt ASML's High NA EUV lithography for mass DRAM production by 2028 and back a move to 12-inch photomasks.
Read original https://semiconductor.samsung.com/news-events/news/samsung-electronics-and-asml-expand-strategic-collaboration-for-next-generation-semiconductor-manufacturingdramlithographymemorysemiconductordramlithographymemorysemiconductor
Sep 6, 20261
- Samsung Semiconductor · Sep 6, 2026Samsung collects SAFE Forum insights on foundry and packaging tech
The resource hub compiles Samsung Foundry sessions and blogs covering 3D chip packaging, process nodes, and semiconductor partnerships.
Read originalsemiconductor.samsung.com/events/safe-forum/insightsSamsung Semiconductor · Sep 6, 2026
Samsung collects SAFE Forum insights on foundry and packaging tech
The resource hub compiles Samsung Foundry sessions and blogs covering 3D chip packaging, process nodes, and semiconductor partnerships.
Read original https://semiconductor.samsung.com/events/safe-forum/insightsno tagsno tags
Sep 4, 20261
- DDN · Sep 4, 2026DDN will showcase AI storage systems at ALL IN Montreal
DDN will highlight its Infinia storage platform designed to keep GPU clusters saturated at the ALL IN summit in Montreal.
Read originalddn.com/lp/events/all-in-canada-2026/book-a-meetingDDN · Sep 4, 2026
DDN will showcase AI storage systems at ALL IN Montreal
DDN will highlight its Infinia storage platform designed to keep GPU clusters saturated at the ALL IN summit in Montreal.
Read original https://ddn.com/lp/events/all-in-canada-2026/book-a-meetingno tagsno tags
Sep 3, 20267
- VAST Data · Sep 3, 2026VAST pitched NVMe-over-TCP remote storage for AI on Kubernetes
VAST says its NVMe/TCP block driver can handle millions of devices per system with automatic host pruning and multi-tenancy for AI workloads.
Read originalvastdata.com/blog/the-case-for-ai-on-kubernetes-remote-storageVAST Data · Sep 3, 2026
VAST pitched NVMe-over-TCP remote storage for AI on Kubernetes
VAST says its NVMe/TCP block driver can handle millions of devices per system with automatic host pruning and multi-tenancy for AI workloads.
Read original https://vastdata.com/blog/the-case-for-ai-on-kubernetes-remote-storageno tagsno tags - NVIDIA · Sep 3, 2026NVIDIA is buying Hugging Face for $12.9 billion
Jensen Huang says the deal covers purchase price plus retention awards and is subject to regulatory approval.
Read originalblogs.nvidia.com/blog/nvidia-to-acquire-hugging-face/NVIDIA · Sep 3, 2026
NVIDIA is buying Hugging Face for $12.9 billion
Jensen Huang says the deal covers purchase price plus retention awards and is subject to regulatory approval.
Read original https://blogs.nvidia.com/blog/nvidia-to-acquire-hugging-face/ainvidiabusinessopen-sourceainvidiabusinessopen-source - Google Cloud HPC · Sep 3, 2026Google Cloud plans a showcase for Lustre, Sycomp, VAST, and Weka storage
The session will compare how DDN Lustre, IBM Storage Scale, VAST Data, and Weka handle high-throughput HPC and AI workloads on Google Cloud.
Read originalsites.google.com/view/advancedcomputingcommunity/high-performance-data-platformsGoogle Cloud HPC · Sep 3, 2026
Google Cloud plans a showcase for Lustre, Sycomp, VAST, and Weka storage
The session will compare how DDN Lustre, IBM Storage Scale, VAST Data, and Weka handle high-throughput HPC and AI workloads on Google Cloud.
Read original https://sites.google.com/view/advancedcomputingcommunity/high-performance-data-platformsstoragelustrecloudhpcstoragelustrecloudhpc - SK hynix · Sep 3, 2026SK hynix showcased HBM4, MRDIMMs, and CXL memory for AI servers
SK hynix displayed server memory products including HBM4, a 128GB MRDIMM, and a liquid-cooled enterprise SSD at Dell Technologies Forum.
Read originalnews.skhynix.com/en/dtf-2026SK hynix · Sep 3, 2026
SK hynix showcased HBM4, MRDIMMs, and CXL memory for AI servers
SK hynix displayed server memory products including HBM4, a 128GB MRDIMM, and a liquid-cooled enterprise SSD at Dell Technologies Forum.
Read original https://news.skhynix.com/en/dtf-2026no tagsno tags - IBM · Sep 3, 2026IBM details Storage Scale System 6000 for AI and HPC clusters
IBM says a single 4U Storage Scale System 6000 reaches up to 330 GB/s throughput and 13 million IOPS with GPUDirect Storage support.
Read originalibm.com/downloads/documents/us-en/10a99803f3afd804IBM · Sep 3, 2026
IBM details Storage Scale System 6000 for AI and HPC clusters
IBM says a single 4U Storage Scale System 6000 reaches up to 330 GB/s throughput and 13 million IOPS with GPUDirect Storage support.
Read original https://ibm.com/downloads/documents/us-en/10a99803f3afd804no tagsno tags - Cornelis Networks · Sep 3, 2026Cornelis Networks lists recent HPC interconnect deployments and partnerships
The vendor newsroom highlights interconnect deployments on supercomputers at LLNL and TACC alongside partnerships with NEC, AMD, and NextSilicon.
Read originalcornelis.com/company/newsroom/storiesCornelis Networks · Sep 3, 2026
Cornelis Networks lists recent HPC interconnect deployments and partnerships
The vendor newsroom highlights interconnect deployments on supercomputers at LLNL and TACC alongside partnerships with NEC, AMD, and NextSilicon.
Read original https://cornelis.com/company/newsroom/storiesno tagsno tags - Dell Technologies · Sep 3, 2026Firebird AI brought an Nvidia Blackwell cluster online in Armenia
The $500 million site uses Dell servers with Nvidia Blackwell GPUs, and the company plans to eventually scale capacity past 400 megawatts.
Read originaldell.com/en-us/blog/firebird-ai-and-dell-advance-sovereign-ai-in-armenia-with-new-ai-clusterDell Technologies · Sep 3, 2026
Firebird AI brought an Nvidia Blackwell cluster online in Armenia
The $500 million site uses Dell servers with Nvidia Blackwell GPUs, and the company plans to eventually scale capacity past 400 megawatts.
Read original https://dell.com/en-us/blog/firebird-ai-and-dell-advance-sovereign-ai-in-armenia-with-new-ai-clusterno tagsno tags
Sep 2, 20263
- SambaNova · Sep 2, 2026SambaNova detailed multi-rack scaling for its SN50 dataflow chip
They project scaling DeepSeek-R1 to 256 RDUs delivers 500 tokens per second while maintaining 45% model bandwidth utilization.
Read originalsambanova.ai/blog/hot-chips-2026-dataflow-at-scaleSambaNova · Sep 2, 2026
SambaNova detailed multi-rack scaling for its SN50 dataflow chip
They project scaling DeepSeek-R1 to 256 RDUs delivers 500 tokens per second while maintaining 45% model bandwidth utilization.
Read original https://sambanova.ai/blog/hot-chips-2026-dataflow-at-scaledataflowarchitectureai-acceleratordataflowarchitectureai-accelerator - HPE / Cray · Sep 2, 2026Oracle tapped HPE Juniper networking for AI data centers
Oracle plans to deploy HPE Juniper routing and switching gear across its cloud data centers to support large-scale AI infrastructure.
Read originalwww.hpe.com/us/en/newsroom/press-release/2026/09/hpe-and-oracle-deepen-networking-collaboration-to-accelerate-gigawatt-scale-ai-infrastructure.htmlHPE / Cray · Sep 2, 2026
Oracle tapped HPE Juniper networking for AI data centers
Oracle plans to deploy HPE Juniper routing and switching gear across its cloud data centers to support large-scale AI infrastructure.
Read original https://www.hpe.com/us/en/newsroom/press-release/2026/09/hpe-and-oracle-deepen-networking-collaboration-to-accelerate-gigawatt-scale-ai-infrastructure.htmlno tagsno tags - Crusoe · Sep 2, 2026Crusoe deploys ON.energy battery systems across 5 GW of AI data centers
Crusoe says the 5 GW battery setup buffers rapid GPU power swings and keeps clusters online during grid voltage drops.
Read originalcrusoe.ai/resources/blog/why-crusoe-chose-on-energys-medium-voltage-ai-ups-solutionCrusoe · Sep 2, 2026
Crusoe deploys ON.energy battery systems across 5 GW of AI data centers
Crusoe says the 5 GW battery setup buffers rapid GPU power swings and keeps clusters online during grid voltage drops.
Read original https://crusoe.ai/resources/blog/why-crusoe-chose-on-energys-medium-voltage-ai-ups-solutionno tagsno tags
Sep 1, 202610
- Arm · Sep 1, 2026Arm explained cache coherency behavior on Neoverse N2
Arm noted that Neoverse N2 upgrades DC IVAC to DC CIVAC and flushes dirty system-level cache lines out to main memory.
Read originalsupport.arm.com/documentation/ka005247/1-0Arm · Sep 1, 2026
Arm explained cache coherency behavior on Neoverse N2
Arm noted that Neoverse N2 upgrades DC IVAC to DC CIVAC and flushes dirty system-level cache lines out to main memory.
Read original https://support.arm.com/documentation/ka005247/1-0architecturecmn-700interconnectneoverse n2architecturecmn-700interconnectneoverse n2 - Intel · Sep 1, 2026Intel detailed Diamond Rapids Xeons and Crescent Island inference GPUs
Intel says Diamond Rapids will pack up to 256 cores and 16 memory channels, while Crescent Island offers 480GB of LPDDR5X in a 350W card.
Read originalintel.com/content/www/us/en/newsroom/news/client-computing/intel-outlines-architectures-for-agentic-ai-at-hot-chips-2026.htmlIntel · Sep 1, 2026
Intel detailed Diamond Rapids Xeons and Crescent Island inference GPUs
Intel says Diamond Rapids will pack up to 256 cores and 16 memory channels, while Crescent Island offers 480GB of LPDDR5X in a 350W card.
Read original https://intel.com/content/www/us/en/newsroom/news/client-computing/intel-outlines-architectures-for-agentic-ai-at-hot-chips-2026.htmlarchitecturecpudiamond rapidshot chipsintelxeonarchitecturecpudiamond rapidshot chipsintelxeon - Samsung Semiconductor · Sep 1, 2026Samsung detailed 3D-stacked zHBM and PIM memory architectures for AI
They claim vertically stacking HBM directly onto processors using copper bonding boosts bandwidth and cuts the energy required for data movement.
Read originalsemiconductor.samsung.com/news-events/tech-blog/memory-innovation-powering-a-new-ai-infrastructure-cycle-ep2Samsung Semiconductor · Sep 1, 2026
Samsung detailed 3D-stacked zHBM and PIM memory architectures for AI
They claim vertically stacking HBM directly onto processors using copper bonding boosts bandwidth and cuts the energy required for data movement.
Read original https://semiconductor.samsung.com/news-events/tech-blog/memory-innovation-powering-a-new-ai-infrastructure-cycle-ep2no tagsno tags - Nebius · Sep 1, 2026Nebius released Aether 3.6 for AI cloud management
The update adds tools for managing and operating production AI infrastructure, the company says.
Read originalnebius.comNebius · Sep 1, 2026
Nebius released Aether 3.6 for AI cloud management
The update adds tools for managing and operating production AI infrastructure, the company says.
Read original https://nebius.comcloudai-infrastructurecloudai-infrastructure - CoreWeave · Sep 1, 2026CoreWeave detailed its Slurm-on-Kubernetes setup for Nvidia clusters
They claim their Slurm on Kubernetes stack reaches up to 96% training goodput on Hopper and improves mean time to failure by about 10x.
Read originalcoreweave.com/blog/what-comes-next-operating-and-evolving-the-production-ai-factoryCoreWeave · Sep 1, 2026
CoreWeave detailed its Slurm-on-Kubernetes setup for Nvidia clusters
They claim their Slurm on Kubernetes stack reaches up to 96% training goodput on Hopper and improves mean time to failure by about 10x.
Read original https://coreweave.com/blog/what-comes-next-operating-and-evolving-the-production-ai-factoryslurmkubernetesgpucloudslurmkubernetesgpucloud - Microsoft Azure HPC · Sep 1, 2026Microsoft previewed Azure Cobalt 200 Arm virtual machines
Microsoft opened preview access for its Cobalt 200 Arm VMs, claiming a 50% performance boost for Linux-based AI workloads.
Read originalazure.microsoft.com/en-us/blog/category/computeMicrosoft Azure HPC · Sep 1, 2026
Microsoft previewed Azure Cobalt 200 Arm virtual machines
Microsoft opened preview access for its Cobalt 200 Arm VMs, claiming a 50% performance boost for Linux-based AI workloads.
Read original https://azure.microsoft.com/en-us/blog/category/computeai infrastructurearmazure cobaltcloud computeai infrastructurearmazure cobaltcloud compute - SK hynix · Sep 1, 2026SK hynix showcased 48GB HBM4 and HBM4E memory at FMS
The company showed 16-layer 48GB HBM4 and 12-layer HBM4E, while proposing High Bandwidth Flash to sit between HBM and SSD tiers.
Read originalnews.skhynix.com/en/fms-2026SK hynix · Sep 1, 2026
SK hynix showcased 48GB HBM4 and HBM4E memory at FMS
The company showed 16-layer 48GB HBM4 and 12-layer HBM4E, while proposing High Bandwidth Flash to sit between HBM and SSD tiers.
Read original https://news.skhynix.com/en/fms-2026cxlhbm3ehbm4hbm4ecxlhbm3ehbm4hbm4e - Arm · Sep 1, 2026Arm detailed 10 MPKI telemetry metrics for Neoverse N1
The updated specification defines formulas for 10 metrics tracking branch mispredictions, TLB walks, and cache misses per thousand instructions.
Read originalsupport.arm.com/documentation/108070/0200/Metrics-by-metric-group-in-Neoverse-N1/MPKI-metrics-for-Neoverse-N1Arm · Sep 1, 2026
Arm detailed 10 MPKI telemetry metrics for Neoverse N1
The updated specification defines formulas for 10 metrics tracking branch mispredictions, TLB walks, and cache misses per thousand instructions.
Read original https://support.arm.com/documentation/108070/0200/Metrics-by-metric-group-in-Neoverse-N1/MPKI-metrics-for-Neoverse-N1benchmarkingneoverse n1profilingtelemetrybenchmarkingneoverse n1profilingtelemetry - Intel · Sep 1, 2026Intel announced Xeon 6+ chips on its 18A process node
They showed 18A-based Xeon 6+ processors and new rack systems pairing Xeon CPUs with SambaNova SN-50 RDUs for inference.
Read originalintel.com/content/www/us/en/newsroom/news/artificial-intelligence/intel-announces-new-ai-innovations-at-computex.htmlIntel · Sep 1, 2026
Intel announced Xeon 6+ chips on its 18A process node
They showed 18A-based Xeon 6+ processors and new rack systems pairing Xeon CPUs with SambaNova SN-50 RDUs for inference.
Read original https://intel.com/content/www/us/en/newsroom/news/artificial-intelligence/intel-announces-new-ai-innovations-at-computex.htmldata centerhardwarexeondata centerhardwarexeon - Gigabyte / Giga Computing · Sep 1, 2026Gigabyte detailed an 8U server for eight AMD MI355X GPUs
The 8U chassis pairs dual AMD EPYC 9005 CPUs with eight MI355X OAM accelerators and twelve 3,000 W power supplies.
Read originalgigabyte.com/Enterprise/GPU-Server/G893-ZX1-AAX4Gigabyte / Giga Computing · Sep 1, 2026
Gigabyte detailed an 8U server for eight AMD MI355X GPUs
The 8U chassis pairs dual AMD EPYC 9005 CPUs with eight MI355X OAM accelerators and twelve 3,000 W power supplies.
Read original https://gigabyte.com/Enterprise/GPU-Server/G893-ZX1-AAX4ai hardwareamd instinctepycgpu serverai hardwareamd instinctepycgpu server
Aug 31, 20263
- Cerebras · Aug 31, 2026Cerebras serves OpenAI GPT-5.6 Sol at 750 tokens per second
They claim pipelining the unquantized model across CS-3 wafer systems hits 750 tokens per second by keeping weights in local SRAM.
Read originalcerebras.ai/blog/how-cerebras-serves-gpt-5-6-sol-at-up-to-750-tokens-per-secondCerebras · Aug 31, 2026
Cerebras serves OpenAI GPT-5.6 Sol at 750 tokens per second
They claim pipelining the unquantized model across CS-3 wafer systems hits 750 tokens per second by keeping weights in local SRAM.
Read original https://cerebras.ai/blog/how-cerebras-serves-gpt-5-6-sol-at-up-to-750-tokens-per-secondbenchmarksinferencellmbenchmarksinferencellm - VAST Data · Aug 31, 2026Glenn Lockwood reviewed why three major HPC and AI system bets stalled
He notes Cori's DataWarp burst buffer saw only 15% utilization despite a 2x I/O boost, and Gordon dropped from 32 planned vSMP nodes to just two.
Read originalvastdata.com/blog/failure-is-a-finding-the-hard-lessons-behind-computings-biggest-betsVAST Data · Aug 31, 2026
Glenn Lockwood reviewed why three major HPC and AI system bets stalled
He notes Cori's DataWarp burst buffer saw only 15% utilization despite a 2x I/O boost, and Gordon dropped from 32 planned vSMP nodes to just two.
Read original https://vastdata.com/blog/failure-is-a-finding-the-hard-lessons-behind-computings-biggest-betsno tagsno tags - Cerebras · Aug 31, 2026Cerebras posted press assets and datasheets for CS-4 and WSE-3 Turbo
The company updated its press kit with photos and datasheets for its CS-3, CS-4, and Wafer-Scale Cluster systems.
Read originalcerebras.ai/company/press-kitCerebras · Aug 31, 2026
Cerebras posted press assets and datasheets for CS-4 and WSE-3 Turbo
The company updated its press kit with photos and datasheets for its CS-3, CS-4, and Wafer-Scale Cluster systems.
Read original https://cerebras.ai/company/press-kitno tagsno tags
