March 29, 2012

Researcher tests performance of diverse HPC architectures

by Ohio Supercomputer Center

Bokhari tests performance of diverse HPC architectures — Saniyah Bokhari, an Ohio State University graduate student, compared performance measures of parallel supercomputers that employ various architectures, testing systems located at the Ohio Supercomputer Center and Pacific Northwest Laboratory: (1) a Cray XMT, 128 proc., 500 MHz, 1-TB; (2) an IBM x3755, 2.4-GHz Opteron,16-core, 64-GB; and (3) an NVIDIA FX 5800 GPU, 1.296 GHz, 240 cores, 4-GB device memory. Credit: Bokhari

Surveying the wide range of parallel system architectures offered in the supercomputer market, an Ohio State University researcher recently sought to establish some side-by-side performance comparisons.

The journal, Concurrency and Computation: Practice and Experience, in February published, "Parallel solution of the subset-sum problem: an empirical study." The paper is based upon a master's thesis written last year by computer science and engineering graduate student Saniyah Bokhari.

"We explore the parallelization of the subset-sum problem on three contemporary but very different architectures, a 128-processor Cray massively multithreaded machine, a 16-processor IBM shared memory machine, and a 240-core NVIDIA graphics processing unit," said Bokhari. "These experiments highlighted the strengths and weaknesses of these architectures in the context of a well-defined combinatorial problem."

Bokhari evaluated the conventional central processing unit architecture of the IBM 1350 Glenn Cluster at the Ohio Supercomputer Center (OSC) and the less-traditional general-purpose graphic processing unit (GPGPU) architecture, available on the same cluster. She also evaluated the multithreaded architecture of a Cray Extreme Multithreading (XMT) supercomputer at the Pacific Northwest National Laboratory's (PNNL) Center for Adaptive Supercomputing Software.

"Ms. Bokhari's work provides valuable insights into matching the best high performance computing architecture with the computational needs of a given research community," noted Ashok Krishnamurthy, interim co-executive director of OSC. "These systems are continually evolving to incorporate new technologies, such as GPUs, in order to achieve new, higher performance measures, and we must understand exactly what each new innovation offers."

Each of the architectures Bokhari tested fall in the area of parallel computing, where multiple processors are used to tackle pieces of complex problems "in parallel." The subset-sum problem she used for her study is an algorithm with known solutions that is solvable in a period of time that is proportional to the number of objects entered, multiplied by the sum of their sizes. Also, she carefully timed the code runs for solving a comprehensive range of problem sizes.

The results from Bokhari's study illustrate that the subset-sum problem can be parallelized well on all three architectures, although for different ranges of problem sizes. The performances of these three machines under varying problem sizes showed the strengths and weaknesses of the three architectures.

Bokhari concluded that the GPU performs well for problems whose tables fit within the limitations of the device memory. Because GPUs typically have memory sizes in the range of 10 gigabytes (GB), such architectures are best for small problems that have table sizes of approximately thirty billion bits.

She found that the IBM x3755 performed very well on medium-sized problems that fit within its 64-GB memory, but had poor scalability as the number of processors increased and was unable to sustain its performance as the problem size increased. The machine tended to saturate for problem with table sizes of 300 billion bits.

The Cray XMT showed very good scaling for large problems and demonstrated sustained performance as the problem size increased, she said. However, the Cray had poor scaling for small problem sizes, performing best with table sizes of a trillion bits or more.

"In conclusion, we can state that the NVIDIA GPGPU is best suited to small problem sizes; the IBM x3755 performs well for medium sizes, and the Cray XMT is the clear choice for large problems," Bokhari said. "For the XMT, we expect to see better performance for large problem sizes, should memory larger than 1 TB become available."

Provided by Ohio Supercomputer Center

Citation: Researcher tests performance of diverse HPC architectures (2012, March 29) retrieved 11 July 2024 from https://phys.org/news/2012-03-diverse-hpc-architectures.html

This document is subject to copyright. Apart from any fair dealing for the purpose of private study or research, no part may be reproduced without the written permission. The content is provided for information purposes only.

Explore further

More chip cores can mean slower supercomputing, simulation shows

0 shares

Feedback to editors

Researchers develop model to study heavy-quark recombination in quark-gluon plasma

27 minutes ago

A new species of extinct crocodile relative rewrites life on the Triassic coastline

11 hours ago

New method achieves tenfold increase in quantum coherence time via destructive interference of correlated noise

11 hours ago

Mars likely had cold and icy past, new study finds

12 hours ago

Study: Nanoparticle vaccines enhance cross-protection against influenza viruses

12 hours ago

New tools are needed to make water affordable, says study

12 hours ago

Researchers demonstrate how to build 'time-traveling' quantum sensors

12 hours ago

Lion with nine lives breaks record with longest swim in predator-infested waters

13 hours ago

New multimode coupler design advances scalable quantum computing

13 hours ago

High-speed electron camera uncovers new 'light-twisting' behavior in ultrathin material

14 hours ago

Load comments (0)

Researcher tests performance of diverse HPC architectures

Researchers develop model to study heavy-quark recombination in quark-gluon plasma

A new species of extinct crocodile relative rewrites life on the Triassic coastline

New method achieves tenfold increase in quantum coherence time via destructive interference of correlated noise

Mars likely had cold and icy past, new study finds

Study: Nanoparticle vaccines enhance cross-protection against influenza viruses

New tools are needed to make water affordable, says study

Researchers demonstrate how to build 'time-traveling' quantum sensors

Lion with nine lives breaks record with longest swim in predator-infested waters

New multimode coupler design advances scalable quantum computing

High-speed electron camera uncovers new 'light-twisting' behavior in ultrathin material

Relevant PhysicsForums posts

Help with some optimization code for Block Matrices.

Is an API Always Necessary for Server-Client Communication?

5 GHz PC WiFi connection Cybersecurity question

I did this POST message configuration damage to my wifi internet, help

Number of Multiplications in the FFT Algorithm

Newbie question about deep learning

More chip cores can mean slower supercomputing, simulation shows

Green 'Oakley Cluster' to double OSC computing power

New NVIDIA Tesla GPUs Reduce Cost Of Supercomputing By A Factor Of 10

Customizing supercomputers from the ground up

NVIDIA Ushers In the Era of Personal Supercomputing

Rice University Selects Dual-Core AMD Opteron Processors To Power New Research Cluster

Hyphens in paper titles harm citation counts and journal impact factors

A big step toward the practical application of 3-D holography with high-performance computers

Combining multiple CCTV images could help catch suspects

Applying deep learning to motion capture with DeepLabCut

Training artificial intelligence with artificial X-rays

New model for large-scale 3-D facial recognition

Medical Xpress

Tech Xplore

Science X

Researcher tests performance of diverse HPC architectures

Researchers develop model to study heavy-quark recombination in quark-gluon plasma

A new species of extinct crocodile relative rewrites life on the Triassic coastline

New method achieves tenfold increase in quantum coherence time via destructive interference of correlated noise

Mars likely had cold and icy past, new study finds

Study: Nanoparticle vaccines enhance cross-protection against influenza viruses

New tools are needed to make water affordable, says study

Researchers demonstrate how to build 'time-traveling' quantum sensors

Lion with nine lives breaks record with longest swim in predator-infested waters

New multimode coupler design advances scalable quantum computing

High-speed electron camera uncovers new 'light-twisting' behavior in ultrathin material

Relevant PhysicsForums posts

Related Stories

More chip cores can mean slower supercomputing, simulation shows

Green 'Oakley Cluster' to double OSC computing power

New NVIDIA Tesla GPUs Reduce Cost Of Supercomputing By A Factor Of 10

Customizing supercomputers from the ground up

NVIDIA Ushers In the Era of Personal Supercomputing

Rice University Selects Dual-Core AMD Opteron Processors To Power New Research Cluster

Recommended for you

Hyphens in paper titles harm citation counts and journal impact factors

A big step toward the practical application of 3-D holography with high-performance computers

Combining multiple CCTV images could help catch suspects

Applying deep learning to motion capture with DeepLabCut

Training artificial intelligence with artificial X-rays

New model for large-scale 3-D facial recognition

Newsletter sign up

Donate and enjoy an ad-free experience