Daily current affairs on DailyCA
DailyCA NotesStudy notes for competitive exams My revisionsRevisions Sign in

ಕಾರ್ಯಕ್ಷಮತೆ ಮಾಪನ, ಸಮಾನಾಂತರ ಪ್ರೊಸೆಸಿಂಗ್ ಮತ್ತು ವಿಶ್ಲೇಷಣೆ | Performance Metrics, Parallel Computing & Advanced Analysis

Advanced 15 min read Updated

Saved in this browser only. Sign in to keep your ticks on every device. See all revisions due

Key points

  • CPU ಎಕ್ಸಿಕ್ಯೂಶನ್ ಸಮಯ = ಇನ್‌ಸ್ಟ್ರಕ್ಷನ್ ಕೌಂಟ್ * CPI * ಕ್ಲಾಕ್ ಸೈಕಲ್ ಸಮಯ | CPU Execution Time equals Instruction Count times CPI times Clock Cycle Time.
  • ಆಮ್ಡಾಲ್ ನಿಯಮದ ಪ್ರಕಾರ ಗರಿಷ್ಠ ವೇಗವರ್ಧನೆಯು ಸೀಕ್ವೆನ್ಷಿಯಲ್ ಭಾಗದಿಂದ ಸೀಮಿತವಾಗಿರುತ್ತದೆ | Amdahl's law defines that speedup is bottlenecked by the sequential fraction.
  • ಪ್ರೊಸೆಸರ್‌ಗಳು ಅನಂತವಾದರೂ ವೇಗವರ್ಧನೆಯ ಮಿತಿ 1/(1-f) ಆಗಿರುತ್ತದೆ | As processing units approach infinity, maximum speedup strictly limits to 1/(1-f).
  • MIPS ಮಾಪನವು ವಿವಿಧ ಆರ್ಕಿಟೆಕ್ಚರ್‌ಗಳ ಹೋಲಿಕೆಗೆ ಸಮಂಜಸವಲ್ಲ | MIPS is unreliable when comparing processors with differing instruction sets.
  • FLOPS ವೈಜ್ಞಾನಿಕ ಮತ್ತು ಗ್ರಾಫಿಕಲ್ ಸಂಸ್ಕರಣೆಯ ಪ್ರಮುಖ ಮಾನದಂಡವಾಗಿದೆ | Floating-point throughput (FLOPS) is the standard metric for compute-intensive tasks.
  • ಕ್ಲಾಕ್ ಆವರ್ತನ ಹೆಚ್ಚಾದಂತೆ ಕ್ಲಾಕ್ ಸೈಕಲ್ ಸಮಯವು ವಿಲೋಮವಾಗಿ ಕಡಿಮೆಯಾಗುತ್ತದೆ | Clock frequency is the mathematical reciprocal of clock cycle duration.
  • ಫ್ಲಿನ್ ಅವರ MIMD ಆರ್ಕಿಟೆಕ್ಚರ್ ಆಧುನಿಕ ಮಲ್ಟಿ-ಕೋರ್ ಪ್ರೊಸೆಸರ್‌ಗಳ ಆಧಾರವಾಗಿದೆ | Flynn's MIMD paradigm represents modern symmetric multiprocessors and multi-core chips.

ಕಂಪ್ಯೂಟರ್ ಕಾರ್ಯಕ್ಷಮತೆ ಮಾಪನ ಮತ್ತು ಕ್ಲಾಕ್ ಸಮೀಕರಣಗಳು | Computer Performance Metrics & Clock Equations

ಕಂಪ್ಯೂಟರ್‌ನ ಕಾರ್ಯಕ್ಷಮತೆಯನ್ನು ಪ್ರಮುಖವಾಗಿ ಪ್ರೋಗ್ರಾಂ ಎಕ್ಸಿಕ್ಯೂಶನ್ ಸಮಯದಿಂದ (CPU Execution Time) ಅಳೆಯಲಾಗುತ್ತದೆ | The definitive measure of computer performance is program execution time, which quantifies the wall-clock time required to complete a given workload.

ಉದಾಹರಣೆ | Example

ಒಂದು ಪ್ರೋಗ್ರಾಂನ CPU ಎಕ್ಸಿಕ್ಯೂಶನ್ ಸಮಯವನ್ನು ಈ ಕೆಳಗಿನ ಸೂತ್ರದಿಂದ ಲೆಕ್ಕಹಾಕಲಾಗುತ್ತದೆ | CPU execution time is given by:

CPU Time = Instruction Count (IC) * Cycles Per Instruction (CPI) * Clock Cycle Time
CPU Time = (Instruction Count * CPI) / Clock Frequency

ಆಮ್ಡಾಲ್ ನಿಯಮ ಮತ್ತು ವೇಗವರ್ಧನೆಯ ಮಿತಿ | Amdahl's Law & Speedup Bounds

ಬಹು-ಪ್ರೊಸೆಸರ್‌ಗಳನ್ನು ಸೇರಿಸಿದಾಗ ಸಿಸ್ಟಮ್ ಎಷ್ಟು ವೇಗವಾಗುತ್ತದೆ ಎಂಬುದನ್ನು ಆಮ್ಡಾಲ್ ನಿಯಮವು ನಿರ್ಧರಿಸುತ್ತದೆ | Amdahl's Law establishes the theoretical maximum speedup possible when using multiple execution cores or parallel resources.

Speedup = 1 / [ (1 - f) + (f / N) ]
ಇಲ್ಲಿ (Where):
f = ಸಮಾನಾಂತರಗೊಳಿಸಬಹುದಾದ ಪ್ರೋಗ್ರಾಂ ಭಾಗ (Parallel fraction of task)
(1 - f) = ಕಡ್ಡಾಯ ಸೀಕ್ವೆನ್ಷಿಯಲ್ ಭಾಗ (Strictly sequential fraction)
N = ಪ್ರೊಸೆಸರ್‌ಗಳ ಅಥವಾ ಕೋರ್‌ಗಳ ಸಂಖ್ಯೆ (Number of processor units)

ಮುಖ್ಯಾಂಶ | Key point

ಪ್ರೋಗ್ರಾಂನ ಶೇಕಡಾ 20 ರಷ್ಟು ಭಾಗವು ಕಡ್ಡಾಯವಾಗಿ ಸೀರಿಯಲ್ ಆಗಿದ್ದರೆ (1 - f = 0.20), ಪ್ರೊಸೆಸರ್‌ಗಳ ಸಂಖ್ಯೆ ಅನಂತವಾದರೂ (N -> infinity), ಗರಿಷ್ಠ ವೇಗವರ್ಧನೆಯು 1 / 0.20 = 5 ಕ್ಕಿಂತ ಹೆಚ್ಚಾಗಲು ಸಾಧ್ಯವಿಲ್ಲ | If 20% of a task is inherently sequential, no matter how many cores are deployed, the theoretical speedup can never exceed 5.

ಕಾರ್ಯಕ್ಷಮತೆ ಮೆಟ್ರಿಕ್‌ಗಳ ಹೋಲಿಕೆ | Comparison of Performance Metrics

ಮೆಟ್ರಿಕ್ | Metric ಪೂರ್ಣ ರೂಪ | Full Form ಮಿತಿಗಳು / ಡೆಫಿನಿಷನ್ | Definition & Limits
MIPS ಮಿಲಿಯನ್ ಇನ್‌ಸ್ಟ್ರಕ್ಷನ್ಸ್ ಪರ್ ಸೆಕೆಂಡ್ (Million Instructions Per Second) ಸೂಚನೆಗಳ ಸಂಕೀರ್ಣತೆಯನ್ನು ನಿರ್ಲಕ್ಷಿಸುತ್ತದೆ (RISC vs CISC ಹೋಲಿಕೆಗೆ ಸೂಕ್ತವಲ್ಲ) | Fails across different ISAs with differing workloads.
MFLOPS / TFLOPS ಫ್ಲೋಟಿಂಗ್ ಪಾಯಿಂಟ್ ಕಾರ್ಯಾಚರಣೆಗಳು (Million/Tera Floating Point Operations) ವೈಜ್ಞಾನಿಕ ಮತ್ತು ಇಂಜಿನಿಯರಿಂಗ್ ಲೆಕ್ಕಾಚಾರಗಳಿಗೆ ಅತ್ಯಂತ ವಿಶ್ವಾಸಾರ್ಹ | Ideal benchmark for scientific workloads and supercomputers.
CPI ಕ್ಲಾಕ್ ಸೈಕಲ್ಸ್ ಪರ್ ಇನ್‌ಸ್ಟ್ರಕ್ಷನ್ (Clock Cycles Per Instruction) ಪ್ರತಿ ಸೂಚನೆಯನ್ನು ಪೂರ್ಣಗೊಳಿಸಲು ಬೇಕಾದ ಸರಾಸರಿ ಕ್ಲಾಕ್ ಚಕ್ರಗಳು | Average number of clock cycles required per instruction.

ಪರೀಕ್ಷಾ ಸಲಹೆ | Exam tip

ಸೂಪರ್‌ಕಂಪ್ಯೂಟರ್‌ಗಳಲ್ಲಿ LINPACK ಬೆಂಚ್‌ಮಾರ್ಕ್ ಬಳಸಿ ರೇಟಿಂಗ್ ನೀಡಲಾಗುತ್ತದೆ. ಟಾಪ್ 500 ಪಟ್ಟಿಯಲ್ಲಿ ಕಾರ್ಯಕ್ಷಮತೆಯನ್ನು Rmax ಮತ್ತು Rpeak (TFLOPS/PFLOPS) ಗಳಲ್ಲಿ ನಮೂದಿಸಲಾಗುತ್ತದೆ | Supercomputer rankings in TOP500 evaluate double-precision matrix solutions via the LINPACK benchmark suite.

Practice questions

Answer all, then check. Explanations appear after checking.

1ಒಂದು ಕಂಪ್ಯೂಟರ್ ಪ್ರೋಗ್ರಾಂ 10^9 ಇನ್‌ಸ್ಟ್ರಕ್ಷನ್‌ಗಳನ್ನು ಹೊಂದಿದೆ. ಪ್ರೊಸೆಸರ್‌ನ ಸರಾಸರಿ CPI (Cycles Per Instruction) 2.0 ಆಗಿದ್ದು, ಕ್ಲಾಕ್ ಆವರ್ತನವು 2 GHz ಆಗಿದೆ. ಈ ಪ್ರೋಗ್ರಾಂ ಅನ್ನು ರನ್ ಮಾಡಲು ತೆಗೆದುಕೊಳ್ಳುವ CPU Execution Time ಎಷ್ಟು? | A program executes 10 to the power 9 instructions on a 2 GHz processor with an average Cycles Per Instruction (CPI) of 2.0. What is the total CPU execution time?
2ಆಮ್ಡಾಲ್ ನಿಯಮದ (Amdahl's Law) ಪ್ರಕಾರ, ಒಂದು ಪ್ರೋಗ್ರಾಂನ ಶೇಕಡಾ 80 ರಷ್ಟು ಭಾಗವನ್ನು ಸಮಾನಾಂತರಗೊಳಿಸಬಹುದು (Parallelizable, f = 0.80). ಪ್ರೊಸೆಸರ್‌ಗಳ ಸಂಖ್ಯೆಯನ್ನು ಅನಂತಕ್ಕೆ (N -> infinity) ಹೆಚ್ಚಿಸಿದರೂ ಗರಿಷ್ಠ ಸಾಧ್ಯವಿರುವ ಸೈದ್ಧಾಂತಿಕ ವೇಗವರ್ಧನೆ (Maximum Theoretical Speedup) ಎಷ್ಟು? | According to Amdahl's Law, if 80% of an application code can be executed in parallel (f = 0.80), what is the maximum theoretical speedup achievable as the number of parallel processing units approaches infinity?
3ಒಂದು ಪ್ರೊಸೆಸರ್ 4 GHz ಕ್ಲಾಕ್ ಆವರ್ತನದಲ್ಲಿ ಕಾರ್ಯನಿರ್ವಹಿಸುತ್ತದೆ. ಅದರ ಒಂದು ಸಿಂಗಲ್ ಕ್ಲಾಕ್ ಸೈಕಲ್ ಸಮಯ (Clock Cycle Duration) ಎಷ್ಟು? | A high-performance processor runs at a sustained clock frequency of 4 GHz. What is the exact physical duration of a single clock cycle?
4ಒಂದು ಕಂಪ್ಯೂಟರ್ ಸಿಸ್ಟಮ್‌ನಲ್ಲಿ 3 ವಿಧದ ಇನ್‌ಸ್ಟ್ರಕ್ಷನ್‌ಗಳಿವೆ: ಕ್ಲಾಸ್ A: 50% ಇನ್‌ಸ್ಟ್ರಕ್ಷನ್‌ಗಳು, CPI = 1 ಕ್ಲಾಸ್ B: 30% ಇನ್‌ಸ್ಟ್ರಕ್ಷನ್‌ಗಳು, CPI = 2 ಕ್ಲಾಸ್ C: 20% ಇನ್‌ಸ್ಟ್ರಕ್ಷನ್‌ಗಳು, CPI = 5 ಈ ವ್ಯವಸ್ಥೆಯ ಸರಾಸರಿ CPI (Average CPI) ಎಷ್ಟು? | A workload consists of three instruction classes: Class A (50% mix, CPI = 1), Class B (30% mix, CPI = 2), and Class C (20% mix, CPI = 5). What is the global average CPI of this architecture?
5ಆಮ್ಡಾಲ್ ನಿಯಮದ ಪ್ರಕಾರ, ಪ್ರೋಗ್ರಾಂನ ಶೇಕಡಾ 60 ರಷ್ಟು ಭಾಗವನ್ನು (f = 0.60) ಸಮಾನಾಂತರವಾಗಿ 4 ಪ್ರೊಸೆಸರ್‌ಗಳನ್ನು ಬಳಸಿ ಚಲಾಯಿಸಿದರೆ, ಒಟ್ಟಾರೆ ಸಿಸ್ಟಮ್ ವೇಗವರ್ಧನೆ (Overall Speedup) ಎಷ್ಟು? | Under Amdahl's Law, if 60% of an algorithm is parallelized across 4 processors (f = 0.6, N = 4), what is the effective overall system speedup?
6ಒಂದು 32-ಬಿಟ್ ಅಡ್ರೆಸ್ ಬಸ್ ಹೊಂದಿರುವ ಮೈಕ್ರೋಪ್ರೊಸೆಸರ್ ನೇರವಾಗಿ ಬೈಟ್-ಅಡ್ರೆಸೇಬಲ್ (Byte-addressable) ಆಗಿ ಗರಿಷ್ಠ ಎಷ್ಟು ಮೆಮೊರಿಯನ್ನು ಆಕ್ಸೆಸ್ ಮಾಡಬಹುದು? | What is the maximum physical memory addressing capacity of a processor architecture equipped with a 32-bit address bus operating in byte-addressable mode?
7ಪ್ರತಿಪಾದನೆ (A): MIPS (Million Instructions Per Second) ಎಂಬುದು ವಿವಿಧ ಇನ್‌ಸ್ಟ್ರಕ್ಷನ್ ಸೆಟ್ ಆರ್ಕಿಟೆಕ್ಚರ್ (ISA) ಹೊಂದಿರುವ ಕಂಪ್ಯೂಟರ್‌ಗಳ ಕಾರ್ಯಕ್ಷಮತೆಯನ್ನು ಹೋಲಿಸಲು ವಿಶ್ವಾಸಾರ್ಹ ಮಾಪನವಲ್ಲ. ಕಾರಣ (R): RISC ಮತ್ತು CISC ಕಂಪ್ಯೂಟರ್‌ಗಳಲ್ಲಿ ಒಂದೇ ಕೆಲಸವನ್ನು ಪೂರ್ಣಗೊಳಿಸಲು ಬೇಕಾಗುವ ಇನ್‌ಸ್ಟ್ರಕ್ಷನ್‌ಗಳ ಸಂಖ್ಯೆ ಮತ್ತು ಪ್ರತಿ ಇನ್‌ಸ್ಟ್ರಕ್ಷನ್‌ನ ಸಂಕೀರ್ಣತೆ ವಿಭಿನ್ನವಾಗಿರುತ್ತದೆ. | Assertion (A): MIPS is fundamentally flawed as an absolute metric when benchmarking hardware platforms across disparate ISAs. Reason (R): Different Instruction Set Architectures (such as RISC versus CISC) demand starkly different instruction counts and cycle complexities to execute identical source code.
8ಒಂದು ಕಂಪ್ಯೂಟರ್ ಸಿಸ್ಟಮ್‌ನ ಕ್ಲಾಕ್ ಆವರ್ತನವು 1 GHz ಆಗಿದ್ದು, ಅದು 2 MIPS ವೇಗವನ್ನು ನೀಡಿದರೆ, ಅದರ ಸರಾಸರಿ CPI ಎಷ್ಟು? | If a computing system operates at 1 GHz clock frequency and achieves an execution throughput of 500 MIPS, what is its average CPI?
9ಆಧುನಿಕ ಸೂಪರ್‌ಕಂಪ್ಯೂಟರ್‌ಗಳಲ್ಲಿ (ಉದಾ: TOP500 ಶ್ರೇಯಾಂಕ) ಸಾಮರ್ಥ್ಯವನ್ನು ಪರೀಕ್ಷಿಸಲು ಬಳಸುವ ಪ್ರಮಾಣಿತ ಗಣಿತೀಯ ಬೆಂಚ್‌ಮಾರ್ಕ್ ಯಾವುದು? | What industry-standard numerical linear algebra benchmark is utilized to rank supercomputers on the international TOP500 list?
10ಒಂದು ವರ್ಕ್‌ಲೋಡ್‌ನ 90% ಭಾಗವನ್ನು ಪ್ಯಾರಲಲ್ ಮಾಡಬಹುದಾಗಿದೆ (f = 0.90). ಆಮ್ಡಾಲ್ ನಿಯಮದ ಪ್ರಕಾರ 9 ಪಟ್ಟು ವೇಗವರ್ಧನೆ (Speedup = 9) ಪಡೆಯಲು ಎಷ್ಟು ಸಮಾನಾಂತರ ಪ್ರೊಸೆಸರ್‌ಗಳು (N) ಅಗತ್ಯವಿದೆ? | If 90% of a computational task is parallelizable (f = 0.90), how many parallel processors (N) must be assigned to achieve an exact system speedup of 9x under Amdahl's Law?
11ಮಾರ್ಕ್ I (Harvard Mark I) ಗಣಕಯಂತ್ರದ ಪ್ರಮುಖ ವಾಸ್ತುಶಿಲ್ಪೀಯ ಮಿತಿ ಯಾವುದಾಗಿತ್ತು? | What was the defining technical characteristic and constraint of the historic 1944 Harvard Mark I electromechanical computer?
12ಒಂದು ಕಂಪ್ಯೂಟರ್‌ನ ಮೆಮೊರಿ ಬಸ್ ಸೈಕಲ್ ಸಮಯ 10 ns ಆಗಿದೆ. ಪ್ರತಿ ಬಸ್ ಸೈಕಲ್‌ನಲ್ಲಿ 64 ಬಿಟ್‌ಗಳ ಡೇಟಾವನ್ನು ವರ್ಗಾಯಿಸಿದರೆ, ಮೆಮೊರಿ ಬಸ್‌ನ ಗರಿಷ್ಠ ಡೇಟಾ ವರ್ಗಾವಣೆ ದರ (Bandwidth) ಎಷ್ಟು? | A system memory bus runs with a cycle duration of 10 nanoseconds. If it transfers 64 bits of data per bus cycle, what is the maximum theoretical bus throughput in MB/s?
13ವಾನ್ ನ್ಯೂಮನ್ ವಾಸ್ತುಶಿಲ್ಪದಲ್ಲಿ ಸೆಲ್ಫ್-ಮಾಡಿಫೈಯಿಂಗ್ ಕೋಡ್ (Self-modifying code) ಸಾಧ್ಯವಾಗಲು ರಚನಾತ್ಮಕ ಕಾರಣವೇನು? | Why is self-modifying code structurally possible within pure Von Neumann architectures but strictly blocked in traditional Harvard machines?
14ಪ್ರಶ್ನೆ: ಗುಸ್ತಫ್ಸನ್ ನಿಯಮವು (Gustafson's Law) ಆಮ್ಡಾಲ್ ನಿಯಮಕ್ಕಿಂತ ಹೇಗೆ ಭಿನ್ನವಾದ ದೃಷ್ಟಿಕೋನವನ್ನು ಪ್ರಸ್ತುತಪಡಿಸುತ್ತದೆ? | How does Gustafson's Law fundamentally counter or expand upon the parallel limits established by Amdahl's Law?
15ಒಂದು ಹೊಸ ಕಂಪ್ಯೂಟರ್ ಆರ್ಕಿಟೆಕ್ಚರ್‌ನಲ್ಲಿ ಸಂಕಲನ ಕಾರ್ಯಾಚರಣೆಯು 40% ಸಮಯವನ್ನು ತೆಗೆದುಕೊಳ್ಳುತ್ತದೆ. ಹೊಸ ALU ವಿನ್ಯಾಸವು ಸಂಕಲನ ಕಾರ್ಯಾಚರಣೆಯನ್ನು 2 ಪಟ್ಟು ವೇಗಗೊಳಿಸಿದರೆ (Speedup of addition = 2), ಒಟ್ಟಾರೆ ಸಿಸ್ಟಮ್ ವೇಗವರ್ಧನೆ ಎಷ್ಟು? | In a computer architecture, addition operations account for 40% of overall runtime. If an enhanced ALU hardware adder accelerates addition operations by a factor of 2x, what is the net overall system speedup?

Finished this topic? Tick it off.

Saved in this browser only. Sign in to keep your ticks on every device. See all revisions due