tutorbin

computer organisation and architecture homework help

Boost your journey with 24/7 access to skilled experts, offering unmatched computer organisation and architecture homework help

tutorbin

Trusted by 1.1 M+ Happy Students

WhatsApp Support

Get Instant
Online Homework Help
via WhatsApp

Get instant homework help from top tutors—just a WhatsApp message away. 24/7 hw help support for all your academic needs!

A
S
M
R
★★★★★
2M+ students trust TutorBin
Your WhatsApp Number
phone
or
⚡ Instant reply
🔒 100% private
👨‍🏫 Top tutors
🌍 All subjects
*Get instant homework help from top tutors—just a WhatsApp message away. 24/7 support for all your academic needs!
2M+ Students Helped24/7 Live SupportExpert TutorsAll Subjects CoveredInstant Response100% ConfidentialTop Rated ServiceMoney-back Guarantee2M+ Students Helped24/7 Live SupportExpert TutorsAll Subjects CoveredInstant Response100% ConfidentialTop Rated ServiceMoney-back Guarantee

Recently Asked computer organisation and architecture Questions

Expert help when you need it
  • Q1:cs281: Introduction to Computer Systems Lab06: The Arithmetic Logic Unit (ALU) 1. The purpose of this laboratory is to: · reinforce our experience in designing and implementing combinational logic circuits, · gain experience in abstraction at the hardware level and assemble larger wholes from smaller parts, and · develop expertise with Logisim. 2. This lab is to be completed individually. You will only submit a copy of your Logisim file; there is no accompa- nying lab report. You will be given a template alu. circ file so that you have a framework from which to build your ALU, and it will help ensure that the external interface (pin definitions and placement) will be consistent with that expected for our grading of your work. 3. We start with a view from outside the ALU circuit you will create. The following picture shows a Logisim main circuit that provides inputs and outputs to a circuit named ALU8. It is the ALU8 circuit that you will design and build in this project. Cout ZF SF OF A ×8 B ×8 ALU8 ×8 F Cin ×3 Op The inputs are shown on the left and bottom of the ALU8 circuit and the outputs are on the right and top of the circuit. As a matter of convention, we use pins on the left of a chip for data inputs, the bottom of the chip for control/select inputs, the right for outputs, and here we use the top of the chip for a set of single-bit outputs that comprise the condition code flags of ALU8. Note that A and B are 8-bit-wide inputs, Op is a 3-bit-wide input, and Cin is a 1-bit-wide input. F is an 8-bit-wide output, and Cout, ZF, SF, and OF are single-bit outputs. 4. The Op determines a particular operation for the combinational logic of the ALU8. So, in general, we can think of the functionality as follows: (F,Cout,ZF,SF,OF)=Op(A,B,Cin) Thus, the set of five outputs are a function, determined by Op, of the other three inputs (A, B, and Cin). 5. In the table below, we give the possible bit patterns for Op, a mneumonic for each operation, and then the com- putation specifying the F output in terms of A, B, and Cin. Since we have been working with C/C++ bitwise operations, we will borrow that notation. Op Mneumonic Semantics 000 add F = B + A + Cin 001 notadd F = B + ~A + Cin 010 and F = A & B 011 xor F=A ^B 100 shiftR F = A >> 1; F7 = Cin 101 shiftL F = A << 1; F0 = Cin 1 Note that the add and notadd also incorporate the value of Cin. Cin is also used as the bit to "fill in" for the shift right (as the most significant bit) and the shift left (as the least significant bit). The reason for this design is to allow this ALU8 to be chained together with three more ALU8 chips to build a 32-bit ALU. Note also that we do not need explicit subtraction, because we can obtain that operation by using the notadd operation and setting Cin to 1, effectively getting the "complement and adding one" semantics we need for two's complement negation. 6. The following table gives the meanings for the four single-bit condition codes. Flag Operations Description ZF all Zero Flag: If the result of the current ALU operation, F, is zero, then ZF is 1. If the result of the operation is not zero, then ZF is 0. SF all Sign Flag: Reflects the most significant bit of the result of the current ALU operation (i.e. F7). When interpreted as an 8-bit two's complement, the sign bit indicates a negative result. add Overflow Flag: This bit is asserted if the current operation caused a two's complement overflow-either positive or negative, and is deasserted otherwise. OF notadd and xor The OF flag is 0 for all four of these other operations. shiftR shiftL add Carry Flag: This bit is 1 if the addition caused a carry-out from the most Cout notadd significant bit position, so an unsigned overflow. shiftR This is the bit "shifted off" from operand A, a.k.a. A0. shiftL This is the bit "shifted off" from operand A, a.k.a. A7. and The Cout flag is 0 for both of these other operations. xor 7. Using the Logisim digital logic design and simulator package, design and implement the above-described ALU8. Start with the ALU. circ file provided on Canvas. Using this template is covered in more detail in the guidance section below. Your implementation may only use 1-bit wide elements, although you are encouraged to build sub circuits that create multi-bit devices from 1-bit wide elements. You may use the built-in multiplexors. 8. This lab will be graded based on a total of 100 points. Of these, 80 of points will come from correct operation of the ALU for all of my test cases. I will be grading for both the output F as well as all four condition codes over a variety of test inputs, so make sure these are implemented correctly per the specification given above. The other 20 points will be based on your circuit design. Some of the grading criteria for the 20 points on these design elements: . Following conventions of inputs on the left, outputs on the right, and being able to read the "flow" of the circuit from left to right. · Abstraction - appropriate use of sub circuits with appropriate pin interfaces and labels so that a problem's decomposition is reflected in the design and hierarchical use of circuits. · Neatness - appropriate "clean" routing of wires and placement of sub circuits to convey an organized and orderly design. I will use a grading circuit that is similar to but larger than the top circuit I have given you. If you move or delete any inputs in your ALU8, then your circuit will not fit in my grading circuit and you will receive a score of 0 for the assignment. It is your responsibility to read all the instructions in this entire handout and follow them carefully. You will submit your alu. circ file for grading. 2 Guidance The guidance given here should help you to order and prioritize your development of the ALU8 circuit. It would be helpful for you to open the ALU. circ file and look at the template structure provided as you read this guidance. 1. When you open the template structure, the current circuit (the one with the magnifying glass icon) should be Manual Top and the design pane should correspond to the top-level picture above. You should make NO changes to this circuit. By extension, this means that you should make no changes to the interface for the ALU8 circuit that would then result in needing to change ManualTop. 2. Notice that, in addition to Manual Top, the circuits named ALU8, Adder8, And8, and Xor8 have been created. If you double-click any of these circuits to see their contents, you will observe that the only elements in each of these are the input pins and the output pins that define the external interface of the circuit. In Logisim, the placement of the pin devices on the design page determine their order and placement when the device is used. Pins on the top (and facing down) of the design panel are on the top of the device when it is used, and the order of the pins from left to right corresponds to the order on the design page. Likewise for pins on the left, right, and bottom of the design page. For each of these already-created circuits, you must keep the top/bottom/left/right positioning of the pins as given to you, but as long as you keep them on the same "edge" and relative position, you can move them to accommodate your circuit design. 3. When you create or edit subcircuits, always do so from the menu at the top or at the left. Never "double click" on an item in the design window. If you double click on an item, you will only modify that one item and the changes will not save globally. If you want to create a new subcircuit, use the menu at the top to select "Project" -> "Add Circuit". 4. You should not add or remove any inputs to your ALU8 circuit. There are three inputs on the left. There is one input on the bottom. There is one output at the right. There are four outputs at the top. You must not add or remove any additional inputs or outputs. If you need additional "input values", then use a constant, not another input. 5. Before you can build an 8-bit wide adder to fill in the body of the Adder8 circuit, you will need to design and build a 1 bit full adder. You can use the design steps from the hardware lab to help this process, with inputs of 1-bit wide A, B, and Cin and 1-bit wide Cout and S as outputs. This should be its own circuit and you should test it well before proceeding. 6. Once the building blocks are in place, focus one at a time on the 8-bit versions of the main building blocks, and make sure a single operation to the ALU works before complicating the design and adding in other high level operations. You should use appropriate width inputs (like 8-bit) and should use splitters, available in the Logisim Wiring folder, to break out the individual single-bit wide wires from the 8-bit aggregate. 7. For some of the condition code bits, the OF in particular, you will want to use our combinational logic design strategy of determining the inputs and building a truth table for the OF output and then designing the circuit to compute the OF. 8. Keep abstraction in mind and build sub-circuits as you go. Retro-fitting sub-circuits is much less helpful in terms of cleaner/simpler circuits and design, and the possibility of leveraging a sub-circuit in multiple places. 3/nSee Answer
  • Q2: COP 5612 Computer System Essentials Module 4: Processor Design (Single Cycle Implementation) and OS Structure 1. (10 pts) Draw the data path block diagram by hand for instruction fetch. Explain in your own words how the hardware works. Note: You should draw by hand the block diagram yourself and explain it in your own words. The intent of this question is that by drawing from scratch will help you understand better. Cut-n-paste from slides or any other source will get you a 0 score on this question. 2. (10 pts) Draw the data path block diagram by hand that executes R-type / Load / Store instructions. Explain in your own words how the hardware works. Note: You should draw by hand the block diagram yourself and explain it in your own words. The intent of this question is that by drawing from scratch will help you understand better. Cut-n-paste from slides or any other source will get you a 0 score on this question. 3. (10 pts) Draw the full data path block diagram by hand for R-type, load, store, and jump instructions. Explain in your own words how the hardware works. Note: You should draw by hand the block diagram yourself and explain it in your own words. The intent of this question is that by drawing from scratch will help you understand better. Cut-n-paste from slides or any other source will get you a 0 score on this question. 4. (7.5 pts) Make a copy of the full data path (Figure 1) provided at the end and highlight the data path when an R-type instruction is executed. 5. (7.5 pts) Make a copy of the full data path (Figure 1) provided at the end and highlight the data path when a load instruction is executed. 6. (7.5 pts) Make a copy of the full data path (Figure 1) provided at the end and highlight the data path when a store instruction is executed. 7. (7.5 pts) Make a copy of the full data path (Figure 1) provided at the end and highlight the data path when a jump instruction is executed. 8. (5 pts) Make a copy of the full data path with control (Figure 2) provided at the end and highlight the data path and show control values when an R-type instruction is executed. 9. (5 pts) Make a copy of the full data path with control (Figure 2) provided at the end and highlight the data path and show control values when a jump instruction is executed. 10. (10 pts) What are six types of system calls? For each type, give an example. 11. (10 pts) With the help of a diagram explain how the following OS is structured. a. (5 pts) MS-DOS b. (5 pts) Traditional Unix 12. (10 pts) Briefly answer the following a. (5 pts) What is a virtual machine? b. (5 pts) List five benefits of a VM. Figure 1 for questions 4, 5, 6, and 7. PC 4 >Add Read address Instruction Instruction memory Read register 1 Read register 2 Write register Write data Figure 2 for questions 8 and 9. Registers Read data 2 RegWrite Read data 1 32 Imm Gen 64 Shift left 1 ALUSrc MUX М 4 Add Sum PCSrc ALU operation Zero ALU ALU result MUX Address MemWrite MemRead Read data Data Write data memory MemtoReg PC 4→>> >Add Read address Instruction [31-0] Instruction memory Instruction [6-0] Instruction [19-15] Instruction [24-20] Instruction [11-7] Control Instruction [31-0] Branch MemRead MemtoReg ALUOP MemWrite ALUSrc RegWrite Read register 1 Read register 2 Write register 32 Read data 1 Read data 2 Write data Registers Imm Gen 64 Instruction [30,14-12] Shift left 1 (OMUX) Add Sum Zero ALU ALU result ALU control EUX - Address Read data Write Data data memory MUXOSee Answer
  • Q3: Assignment 4. Input: D0, n x n matrix with 0 on diagonal, positive values other places Output: D, n x n matrix for k starting from 0 through n-1 for i starting from 0 through n-1 for j starting from 0 through n-1 D[i][j] min{ D0[i][j], D0 [i] [k] + D0 [k] [j] } end for j end for i for i starting from 0 through n-1 for j starting from 0 through n-1 D0 [i] [j] = end for j end for i end for k D[i][j] Write an MPI code that parallelizes the pseudo code given above by parallelizing the inner for-i and for-j loops in a way that each process can allocate memories only for a sub-matrix of D and a sub-matrix of DO of size n/√P × n/√P, where P is the number of processes. You may allocate up to 4 more 1-D arrays of size n/√P. You may assign processes as in the figure below to make the parallelization easier, where the number in each small square is the process id, i.e. the rank of the process. You may assume n is divisible by √P and P is the square of a positive integer. The communication cost of your parallel algorithm should be bounded by 0 (log2 VP) √P for each k-loop iteration. 0 √P √P-1 √P+1 2√√√P-1 ?? ??+1 P-1See Answer
  • Q4: COSC 2425: Computer Organization and Architecture Homework 4 (Chapter 4) Due: April 18th, 2024, 11:59 PM Ques # Points Q1 20 Q2 20 Q3 20 Total Grade /60 Please use the solution sheet to submit your answers. 1. Submissions are only allowed in the format provided in the solution sheet. 2. The solution sheet is a word document (COA_HW4_Solution_Sheet.docx). 3. Download the solution sheet. The sheet has a table with question numbers. 4. Edit the word document by filling in your solution in the corresponding boxes in the table. 5. Points are allocated both for final answer and your work/formula. 6. Please show your work clearly and how you arrived at the solution. Pipelining Question 1: Consider an LEGv8 processor with pipelined implementation consisting of five stages: 1. IF: Instruction Fetch 2. ID: Instruction Decode 3. EXE: Execution 4. MEM: Memory access 5. WB: Write Back Write operations occur in the first half of the clock cycle and reads occur in the second half of the clock cycle. The table below shows the pipelining diagram for three instructions. Inst/CC 1 2 3 4 5 6 7 8 Inst. 1 IF ID EXE MEM WB Stall Inst. 2 Inst. 3 IF ID EXE MEM WB IF ID EXE MEM WB Where CC stands for clock cycle, stalls are shown as empty rows, and forwarding is indicated between associated stages using an arrow. LDUR X0, [X0, #0] ADD X1, X0, XO SUB X2, X1, X2 SUB X3, X2, X3 ADD X4, X3, X2 a. First assume forwarding is not available, show a similar table for the following sequence of instructions and indicate stalls with empty rows. How many cycles are needed to execute the above sequence of instructions? b. This time, assume that forwarding is available. Show a similar table for the following sequence of instructions and indicate stalls with empty rows and forwarding with arrows. How many cycles are needed to execute the above sequence of instructions? Data Hazards Question 2: Consider the following loop 1. LDUR X0, [X11, #0] 2. LDUR X1, [X10, #4] 3. SUB X3, XO, XO 4. ADD X3, X1, X3 5. ADD X4, X3, X1 1. List the true data dependencies (Read-after-write) in the above code. Use the number next to the instruction to identify instructions. For example, If the Instruction with line number x has a dependency on instruction with line number y. Then state that "Instruction x on instruction y". Note: List all dependencies, respective of whether they cause stalls in the pipeline or not. 2. Assume 5-stage pipeline with no forwarding, and each stage takes 1 cycle. a. Show the pipeline diagram for the instruction sequence. b. Assuming that the processor stalls on a hazard. How many times does the processor stall? How long is each stall (in cycles)? What is the execution time (in cycles) for the whole program? c. Assume the 5-stage pipeline with full forwarding. Show the pipeline diagram with stalls if needed. What is the execution time (in cycles) for the whole program? Branch Prediction Question 3: In the context of control hazards, branch prediction strategies are used to improve performance. Given a branch, the following example shows the branch prediction table for some strategy and the outcome of the branch. Predictor value of 1 indicates the prediction that the branch will be taken and 0 indicates the prediction that the branch will not be taken. Value of register Branch predictor for Branch b Prediction (T/NT) Actual outcome Misprediction? of branch bl (yes/no) 1 0 NT NT No 2345 2 0 NT T Yes 3 1 T NT Yes 4 0 NT NT No 5 0 NT NT No Branch predictor accuracy is defined as the percentage of predictions that are correct. In the example above, the prediction is correct 3 out of 5 times (shown in bold). The accuracy can be computed as shown below: Accuracy = 3/5 = 0.6 Let's assume that A is an array stored in the main memory with values [6, 4, 3, 7, 1, 4, 9] associated with array indices 0 through 6, respectively. Furthermore, assume that the base address for array variable A is associated with register XO, and that of i and k are associated with X1, X2, respectively. Consider the following LEGv8 code. AND X1, X1, XZR Loop: LSL X3, X1, #3 ADD X3, X0, X3 LDUR X4, [X3, #0] SUBI X5, X4, #4 CBNZ X5, Cond ADDI X2, X2, #1 Cond: SUBI X6, X1, #7 // Branch B1 CBZ X6, Exit ADDI X1, X1, #1 B Loop Exit: Consider the following strategies and show branch prediction strategies tables and accuracies for branch B1 for each of the strategy A. Strategy 1 → Branch is always taken: In this, the strategy is to assume that the branch will always be taken. For this strategy show the branch prediction table, comparing it with the actual branch outcome and compute the branch predictor accuracy. B. Strategy 2 → Branch is always taken: In this, the strategy is to assume that the branch will never be taken. For this strategy show the branch prediction table, comparing it with the actual branch outcome and compute the branch predictor accuracy. C. Strategy 3 → 1-bit branch predictor: Consider a 1-bit prediction strategy as discussed in class. If the initial value for the predictor is 0, then for this strategy show the branch prediction table and compute the branch predictor accuracy. D. Strategy 4 → 2-bit branch predictor: Consider a 2-bit prediction strategy as discussed in class. If the initial value for the predictor is 01, then for this strategy show the branch prediction table and compute the branch predictor accuracy./nSee Answer
  • Q5: uploaded previously done abstract and the ieee paper that should be implemented Submit code with comments + output screenshot + 2-page report on how code is implemented/nProject Implementation 1. For the project which involves implementations, either demo or presentation (not both) can be performed to discuss the details. The option is left to the individual students. A. For demo a maximum 8 minutes time is allocated for a specific student. The student needs to bring laptop, portable platform, boards, etc. to the class. B. For presentation, please prepare approximately 10 slides to cover a 8 minutes presentation. You should upload the presentations using this link. This will make it easier for you to download them in the class PC and use. 2. Project final report submission. The report can be a 10 to 20 pages writeup. Please provide information of the project as follows: (1) It should have a title. (2) What is the project all about? (3) How much you have implemented already in this project? (4) What is left to be implemented that you could not do? (5) What are the difficulties and challenges you faced for this project? (6) Design and simulation figures and tables. (7) Formal references. Note: (1) The text and figures to be used should be your own. Draw the figures that you want to use. Write in your own words. (2) Use latex for your writing as it is considered better for technical writing./n IEEE TRANSACTIONS ON PARALLEL AND DISTRIBUTED SYSTEMS, VOL. 23, NO. 8, AUGUST 2012 A Survey of Parallel Programming Models and Tools in the Multi and Many-Core Era Javier Diaz, Camelia Muñoz-Caro, and Alfonso Niño Abstract—In this work, we present a survey of the different parallel programming models and tools available today with special consideration to their suitability for high-performance computing. Thus, we review the shared and distributed memory approaches, as well as the current heterogeneous parallel programming model. In addition, we analyze how the partitioned global address space (PGAS) and hybrid parallel programming models are used to combine the advantages of shared and distributed memory systems. The work is completed by considering languages with specific parallel support and the distributed programming paradigm. In all cases, we present characteristics, strengths, and weaknesses. The study shows that the availability of multi-core CPUs has given new impulse to the shared memory parallel programming approach. In addition, we find that hybrid parallel programming is the current way of harnessing the capabilities of computer clusters with multi-core nodes. On the other hand, heterogeneous programming is found to be an increasingly popular paradigm, as a consequence of the availability of multi-core CPUs+GPUs systems. The use of open industry standards like OpenMP, MPI, or OpenCL, as opposed to proprietary solutions, seems to be the way to uniformize and extend the use of parallel programming models. Index Terms-Parallelism and concurrency, distributed programming, heterogeneous (hybrid) systems. 1369 1 INTRODUCTION MICROPROCESSORS bormance increases and cost reducing the of the software applications are developed followy tions in computer applications for more than two decades. However, this process reached a limit around 2003 due to heat dissipation and energy consumption issues [1]. These problems have limited the increase of CPU clock frequen- cies and the number of tasks that can be performed within each clock period. The solution adopted by processor developers was to switch to a model where the micro- processor has multiple processing units known as cores [2]. Nowadays, we can speak of two approaches [2]. The first, multi-core approach, integrates a few cores (currently between two and ten) into a single microprocessor, seeking to keep the execution speed of sequential programs. Actual laptops and desktops incorporate this kind of processor. The second, many-core approach uses a large number of cores (currently as many as several hundred) and is specially oriented to the execution throughput of parallel programs. This approach is exemplified by the Graphical Processing Units (GPUs) available today. Thus, parallel computers are not longer expensive and elitist devices, but commodity machines we find everywhere. Clearly, this change of paradigm has had (and will have) a huge impact on the software developing community [3]. • J. Diaz is with the Pervasive Technology Institute, Indiana University, 2719 East Tenth Street, Bloomington, IN 47408. E-mail: javidiaz@indiana.edu. implemented on traditional single-core microprocessors. Therefore, each new, more efficient, generation of single- core processors translates into a performance increase of the available sequential applications. However, the current stalling of clock frequencies prevents further performance improvements. In this sense, it has been said that "sequential programming is dead" [4], [5]. Thus, in the present scenario we cannot rely on more efficient cores to improve performance but in the appropriate coordinate use of several cores, i.e., in concurrency. So, the applications that can benefit from performance increases with each generation of new multi-core and many-core processors are the parallel ones. This new interest in parallel program development has been called the “concurrency revolution” [3]. Therefore, parallel programming, once almost relegated to the High Performance Computing community (HPC), is taken a new star role on the stage. Parallel computing can increase the applications perfor- mance by executing them on multiple processors. Unfortu- nately, the scaling of application performance has not matched the scaling of peak speed, and the programming burden continues to be important. This is particularly problematic because the vision of seamless scalability needs the applications to scale automatically with the number of processors. However, for this to happen, the applications have to be programmed to exploit parallelism in the most efficient way. Thus, the responsibility for achieving the vision of scalable parallelism falls on the applications developer [6]. In this sense, there are two main approaches to parallelize applications: autoparallelization and parallel programming [7]. They differ in the achievable application performance and ease of parallelization. In the first case, the sequential programs are automatically parallelized using Authorized licensed use limited to: University of North Texas. Downloaded on February 19,2024 at 01:11:54 UTC from IEEE Xplore. Restrictions apply. • C. Muñoz-Caro and A. Niño are with the Grupo SciCom, Departamento de Tecnologias y Sistemas de Informacion, Escuela Superior de Informática, Universidad de Castilla-La Mancha, Paseo de la Universidad 4, 13004 Ciudad Real, Spain. E-mail: {camelia.munoz, alfonso.nino}@uclm.es. Manuscript received 22 Apr. 2011; revised 7 Dec. 2011; accepted 8 Dec. 2011; published online 28 Dec. 2011. Recommended for acceptance by H. Jiang. For information on obtaining reprints of this article, please send e-mail to: tpds@computer.org, and reference IEEECS Log Number TPDS-2011-04-0242. Digital Object Identifier no. 10.1109/TPDS.2011.308. 1045-9219/12/$31.00 © 2012 IEEE Published by the IEEE Computer Society 1370 IEEE TRANSACTIONS ON PARALLEL AND DISTRIBUTED SYSTEMS, VOL. 23, NO. 8, AUGUST 2012 ILP (instruction level parallelism) or parallel compilers. Thus, the main advantage is that existing applications just need to be recompiled with a parallel compiler, without modifications. However, due to the complexity of auto- matically transforming sequential algorithms into parallel ones, the amount of parallelism reached using this approach is low. On the other hand, in the parallel programming approach, the applications are specifically developed to exploit parallelism. Therefore, developing a parallel appli- cation involves the partitioning of the workload into tasks, and the mapping of the tasks into workers (i.e., the computers where the tasks will be processed). In general, parallel programming obtains a higher performance than autoparallelization but at the expense of more paralleliza- tion efforts. Fortunately, there are some typical kinds of parallelism in computer programs such as task, data, recursive, and pipelined parallelism [8], [9], [10]. In addition, much literature is available about the suitability of algorithms for parallel execution [11], [12] and about the design of parallel programs [10], [13], [14], [15]. From the design point of view, different patterns for exploiting parallelism have been proposed [8], [10]. A pattern is a strategy for solving recurring problems in a given field. In addition, the patterns can be organized as part of a pattern language, allowing the user to use the patterns to build complex systems. This approach applied to parallel pro- gramming is presented in [10]. Here, the pattern language is organized in four design spaces or phases: finding con- currency, algorithm structure, supporting structures, and implementation mechanisms. A total of 19 design patterns are recognized and organized around the first three phases. In particular, four patterns corresponding to the supporting structures phase can be related to the different parallel programming models [10]. These are: Single Program Multiple data (SPMD, where the same program is executed several times with different data), Master/Worker (where a master process sets up a pool of worker processes and a bag of tasks), loop parallelism (where different iterations of one or more loops are executed concurrently), and fork/join (where a main process forks off several other processes that execute concurrently until they finally join in a single process again). Parallel systems, or architectures, fall into two broad categories: shared memory and distributed memory [8]. In shared memory architectures we have a single memory address space accessible to all the processors. Shared memory machines have existed for a long time in the servers and high-end workstations segment. However, at present, common desktop machines fall into this category since in multi-core processors all the cores share the main memory. On the other hand, in distributed memory architectures there is not global address space. Each processor owns its own memory. This is a popular architectural model encountered in networked or distributed environments such as clusters or Grids of computers. Of course, hybrid shared- distributed memory systems can be built. distributed memory systems. The largest and fastest computers today employ both shared and distributed memory architectures. This provides flexibility when tuning the parallelism in the programs to generate max- imum efficiency and an appropriate balance of the computational and communication loads. In addition, the availability of General Purpose computation on GPUs (GPGPUs) in actual multi-core systems has lead to the Heterogeneous Parallel Programming (HPP) model. HPP seeks to harness the capabilities of multi-core CPUs and many-core GPUs. Accordingly to all theses hybrid archi- tectures, different parallel programming models can be mixed in what is called hybrid parallel programming. A wise implementation of hybrid parallel programs can generate massive speedups in the otherwise pure MPI or pure OpenMP implementations [18]. The same can be applied to hybrid programming involving GPUs and distributed architectures [19], [20]. In this paper, we review the parallel programming models with especial consideration of their suitability for High Performance Computing applications. In addition, we con- sider the associated programming tools. Thus, in Section 2 we present a classification of parallel programming models in use today. Sections 3 to 8 review the different models presented in Section 2. Finally, in Section 9 we collect the conclusions of the work. 2 CLASSIFICATION OF PARALLEL PROGRAMMING MODELS Strictly speaking, a parallel programming model is an Therefore, it is not tied to any specific machine type. abstraction of the computer system architecture [10]. However, there are many possible models for parallel computing because of the different ways several processors can be put together to build a parallel system. In addition, separating the model from its actual implementation is often difficult. Parallel programming models and its associated implementations, i.e., the parallel programming environments defined by Mattson et al. [10], are over- whelming. However, in the late 1990s two approaches become predominant in the HPC parallel programming landscape: OpenMP for shared memory and MPI for distributed memory [10]. This allows us to define the classical or pure parallel models. In addition, the new processor architectures, multi-core CPUs and many-core GPUs, have produced heterogeneous parallel programming models. Also, the simulation of a global memory space in a distributed environment leads to the Partitioned Global Address Space (PGAS) model. Finally, the architectures available today allow definition of hybrid, shared-distributed memory + GPU, models. The parallel computing landscape would not be complete without considering the languages with parallel support and the distributed programming model. All these topics are presented in the next sections. 3 PURE PARALLEL PROGRAMMING MODELS The conventional parallel programming practice in- volves a pure shared memory model [8], usually using the Here, we consider parallel programming models using a OpenMP API [16], in shared memory architectures, or a pure shared or distributed memory approach. As such, we pure message passing model [8], using the MPI API [17], on consider the threads, shared memory OpenMP, and Authorized licensed use limited to: University of North Texas. Downloaded on February 19,2024 at 01:11:54 UTC from IEEE Xplore. Restrictions apply. DIAZ ET AL.: A SURVEY OF PARALLEL PROGRAMMING MODELS AND TOOLS IN THE MULTI AND MANY-CORE ERA Programming Implementation: TABLE 1 Pure Parallel Programming Models Implementations Pthreads Threads OpenMP Shared Memory MPI Message Passing Model System Architecture Communication Model Granularity Synchronization Shared memory Shared memory Shared Address Shared Address Course or Fine Explicit Fine Implicit Implementation WebPage Library a) Compiler b) Distributed and Shared memory Message Passing or Shared Address Course or Fine Implicit or Explicit Library 1371 distributed memory message passing models. Table 1 collects the characteristics of the usual implementations of these models. 3.1 POSIX Threads In this model, we have several concurrent execution paths (the threads) that can be controlled independently. A thread is a lightweight process having its own program counter and execution stack [9]. The model is very flexible, but low level, and is usually associated with shared memory and operating systems. In 1995 a standard was released [21]: the POSIX.1c, Threads extensions (IEEE Std 1003.1c-1995), or as it is usually called Pthreads. The Pthreads, or Portable Operating System Interface (POSIX) Threads, is a set of C programming language types and procedure calls [7], [22], [23]. Pthreads is implemented as a header (pthread.h) and a library for creating and manipulating each thread. The Pthreads library provides functions for creating and destroying threads and for coordinating thread activities via constructs designed to ensure exclusive access to selected memory locations (locks and condition variables). This model is especially appro- priate for the fork/join parallel programming pattern [10]. In the POSIX model, the dynamically allocated heap memory, and obviously the global variables, is shared by the threads. This can cause programming difficulties. Often, one needs a variable that is global to the routines called within a thread but that is not shared between threads. A set of Pthreads functions is used to manipulate thread local storage to address these requirements. Moreover, when multiple threads access the shared data, programmers have to be aware of race conditions and deadlocks. To protect critical section, i.e., the portion of code where only one thread must reach shared data, Pthreads provides mutex (mutual exclusion) and semaphores [24]. Mutex permits only one thread to enter a critical section at a time, whereas semaphores allow several threads to enter a critical section. of threads is not related to the number of processors available. These characteristics make Pthreads programs not easily scalable to a large number of processors [6]. For all these reasons, the explicitly-managed threads model is not well suited for the development of HPC applications. 3.2 Shared Memory OpenMP Strictly speaking, this is also a multithreaded model, as the previous one. However, here we refer to a shared memory parallel programming model that is task oriented and works at a higher abstraction level than threads. This model is in practice inseparable from its practical implementation: OpenMP. OpenMP [25] is a shared memory application program- ming interface (API) whose aim is to ease shared memory parallel programming. The OpenMP multithreading inter- face [16] is specifically designed to support HPC programs. It is also portable across shared memory architectures. OpenMP differs from Pthreads in several significant ways. While Pthreads is purely implemented as a library, OpenMP is implemented as a combination of a set of compiler directives, pragmas, and a runtime providing both manage- ment of the thread pool and a set of library routines. These directives instruct the compiler to create threads, perform synchronization operations, and manage shared memory. Therefore, OpenMP does require specialized compiler sup- port to understand and process these directives. At present, an increasing number of OpenMP versions for Fortran, C, and C++ are available in free and proprietary compilers, see Appendix 1, which can be found on the Computer Society Digital Library at http://doi.ieeecomputersociety.org/ 10.1109/TPDS.2011.308. In OpenMP the use of threads is highly structured because it was designed specifically for parallel applica- tions. In particular, the switch between sequential and parallel sections of code follows the fork/join model [9]. This is a block-structured approach for introducing con- currency. A single thread of control splits into some number of independent threads (the fork). When all the threads have completed the execution of their specified tasks, they resume the sequential execution (the join). A fork/join block corresponds to a parallel region, which is defined using the PARALLEL and END PARALLEL directives. In general, Pthreads is not recommended as a general- purpose parallel program development technology. While it has its place in specialized situations, and in the hands of expert programmers, the unstructured nature of Pthreads constructs makes difficult the development of correct and maintainable programs. In addition, recall that the number Authorized licensed use limited to: University of North Texas. Downloaded on February 19,2024 at 01:11:54 UTC from IEEE Xplore. Restrictions apply. 1372 IEEE TRANSACTIONS ON PARALLEL AND DISTRIBUTED SYSTEMS, VOL. 23, NO. 8, AUGUST 2012 The parallel region enables a single task to be replicated across a set of threads. However, in parallel programs is very common the distribution of different tasks across a set of threads, such as parallel iterations over the index set of a loop. Thus, there is a set of directives enabling each thread to execute a different task. This procedure is called worksharing. Therefore, OpenMP is specially suited for the loop parallel program structure pattern, although the SPMD and fork/join patterns also benefit from this programming environment [10]. OpenMP provides application-oriented synchronization primitives, which make easier to write parallel programs. By including these primitives as basic OpenMP operations, it is possible to generate efficient code more easily than, for instance, using Pthreads and working in terms of mutex and condition variables. In May 2008 the OpenMP 3.0 version was released [26]. The major change in this version was the support for explicit tasks. Explicit tasks ease the parallelization of applications where units of work are generated dynami- cally, as in recursive structures or in while loops. This new characteristic is very powerful. By supporting while loops and other iterative control structures, it is possible to handle graph algorithms and dynamic data structures, for instance. The characteristics of OpenMP allow for a high abstrac- tion level, making it well suited for developing HPC applications in shared memory systems. The pragma directives make easy to obtain concurrent code from serial code. In addition, the existence of specific directives eases to parallelize loop-based code. However, the high cost of traditional multiprocessor machines prevented the wide- spread use of OpenMP. Nevertheless, the ubiquitous availability of multi-core processors has renewed the interest for this parallel programming model. 3.3 Message Passing Message Passing is a parallel programming model where communication between processes is done by interchanging messages. This is a natural model for a distributed memory system, where communication cannot be achieved by sharing variables. There are more or less pure realizations of this model such as Aggregate Remote Memory Copy Interface (ARMCI), which allows a programming approach between message passing and shared memory. ARMCI is detailed later in Section 5.2.1. However, over time, a standard has evolved and dominated for this model: the Message Passing Interface (MPI). different nodes of computational Grids implementing well- established middlewares such as Globus (the de facto standard, see Section 8.1 later) [30]. MPI addresses the message-passing model [6], [27], [28]. In this model, the processes executed in parallel have separate memory address spaces. Communication occurs when part of the address space of one process is copied into the address space of another process. This operation is cooperative and occurs only when the first process executes a send operation and the second process executes a receive operation. In MPI, the workload partitioning and task mapping have to be done by the programmer, similarly to Pthreads. Programmers have to manage what tasks are to be computed by each process. Communication models in MPI comprise point-to-point, collective, one-sided, and parallel I/O operations. Point-to-point operations such as the "MPI_Send"/"MPI_Recv" pair facilitate communications between processes. Collective operations such as “MPI_ Bcast" ease communications involving more than two processes. Regular MPI send/receive communication uses a two-sided model. This means that matching operations by sender and receiver are required. Therefore, some amount of synchronization is needed to manage the matching of sends and receives, and the associated buffer space, of messages. However, starting from MPI-2 [31], one-sided communica- tions are possible. Here, no sender-receiver matching is needed. Thus, one-sided communication decouples data transfer from synchronization. One-sided communication allows remote memory access. Three communication calls are provided: "MPI_Put” (remote write), “MPI_Get” (re- mote read), and “MPI_Accumulate" (remote update). Finally, parallel I/O is a major component of MPI-2, providing access to external devices exploiting data types and communicators [28]. On the other hand, with Symmetric Multiprocessing (SMP) machines being commonly available, and multi-core processors becoming the norm, a programming model to be considered is a mixture of message passing and multi- threading. In this model, user programs consist of one or more MPI processes on each SMP node or multi-core processor, with each MPI process itself comprising multiple threads. The MPI-2 Standard [31] has clearly defined the interaction between MPI and user created threads in an MPI program. This specification was written with the goal of allowing users to write multithreaded MPI programs easily. Thus, MPI supports four "levels" of thread safety that a user must explicitly select: • MPI THREAD SINGLE. A process has only one thread of execution. MPI THREAD FUNNELED. A process may be multithreaded, but only the thread that initialized MPI can make MPI calls. MPI THREAD SERIALIZED. A process may be multithreaded, but only one thread at a time can make MPI calls. MPI is a specification for message passing operations [6], [27], [28], [29]. MPI is a library, not a language. It specifies the names, calling sequences, and results of the subroutines or functions to be called from Fortran, C or C++ programs. Thus, the programs can be compiled with ordinary compilers but must be linked with the MPI library. MPI is currently the de facto standard for HPC applications on distributed architectures. By its nature it favors the SPMD and, to a lesser extent, the Master/Worker program structure patterns [10]. Appendix 2, available in the online supplemental material, collects some well-known MPI implementations. It is interesting to note that MPICH-G2 and GridMPI are MPI implementations for computational Grid environments. Thus, MPI applications can be run on Authorized licensed use limited to: University of North Texas. Downloaded on February 19,2024 at 01:11:54 UTC from IEEE Xplore. Restrictions apply. MPI THREAD MULTIPLE. A process may be multithreaded and multiple threads can call MPI functions simultaneously. Further details about thread safety is provided in [31]. In addition, in [32] the authors analyze and discuss critical issues of thread-safe MPI implementations. DIAZ ET AL.: A SURVEY OF PARALLEL PROGRAMMING MODELS AND TOOLS IN THE MULTI AND MANY-CORE ERA In summary, MPI is well suited for applications where portability, both in space (across different systems existing now) and in time (across generations of computers), is important. MPI is also an excellent choice for task-parallel computations and for applications where the data structures are dynamic, such as unstructured mesh computations. Over the last two decades (the computer cluster era), message passing, and specifically MPI, has become the HPC standard approach. Thus, most of the current scientific code allows for parallel execution under the message passing model. Examples are: the molecular electronic structure codes NWChem [33] and Gamess [34], or mpiBLAST [35] the parallel version of the Basic Local Alignment Search Tool (BLAST) used to find regions of local similarity between nucleotide or protein sequences. Message passing (colloquially understood as MPI) is so tied to HPC and scientific computing that, at present, in many scientific fields HPC is synonymous of MPI programming. 4 HETEROGENEOUS PARALLEL PROGRAMMING MODELS In the beginning of 2001 NVIDIA introduced the first programmable GPU: GeForce3. Later, in 2003 the Siggraph/ Eurographics Graphics Hardware workshop, held in San Diego, showed a shift from graphics to nongraphics applications of the GPUs [36]. Thus, the GPGPU concept was born. Today, it is possible to have, in a single system, one or more host CPUs and one or more GPUs. In this sense, we can speak of heterogeneous systems. Therefore, a programming model oriented toward these systems has appeared. The heterogeneous model is foreseeable to become a mainstream approach due to the microprocessors industry interest in the development of Accelerated Processing Units (APUs). An APU integrates the CPU (multi-core) and a GPU on the same die. This design allows for a better data transfer rate and lower power consump- tion. AMD Fusion [37] and Intel Sandy Bridge [38] APUs are examples of this tendency. In the first CPU+GPU systems, languages as Brook [39] or Cg [40] were used. However, NVIDIA has popularized (Compute Unified Device Architecture) CUDA [41] as the primary model and language to program their GPUs. More recently, the industry has worked together on the Open Computing Language (OpenCL) standard [42] as a common model for heterogeneous programming. In addition, differ- ent proprietary solutions, such as Microsoft's DirectCompute [43] or Intel's Array Building Blocks (ArBB) [44], are available. This section reviews these approaches. 4.1 CUDA CUDA is a parallel programming model developed by NVIDIA [41]. The CUDA project started at 2006 with the first CUDA SDK released in early 2007. The CUDA model is designed to develop applications scaling transparently with the increasing number of processor cores provided by the GPUs [1], [45]. CUDA provides a software environment that allows developers to use C as high-level programming language. In addition, other languages bindings or application programming interfaces are Grid (Device) Block (Work-group) Registers (Private Registers Registers Registers (Private (Private (Private Memory) Memory) Memory) Memory) Thread (Work-item) Thread (Work-item) Thread (Work-item) Shared Memory (Local Memory) Thread (Work-item) Shared Memory (Local Memory) Global/Constant Memory Host Fig. 1. CUDA (OpenCL) architecture and memory model. 1373 supported; see Appendix 3, available in the online supplemental material. For CUDA, a parallel system consists of a host (i.e., CPU) and a computation resource or device (i.e., GPU). The computation of tasks is done in the GPU by a set of threads running in parallel. The GPU threads architecture consists in a two-level hierarchy, namely the block and the grid, see Fig. 1. The block is a set of tightly coupled threads, each identified by a thread ID. On the other hand, the grid is a set of loosely coupled blocks with similar size and dimension. There is no synchronization at all between the blocks, and an entire grid is handled by a single GPU. The GPU is organized as a collection of multiprocessors, with each multiprocessor responsible for handling one or more blocks in a grid. A block is never divided across multiple multiprocessors. Threads within a block can cooperate by sharing data through some shared memory, and by synchronizing their execution to coordinate memory ac- cesses. More detailed information can be found in [41], [46]. Moreover, there is a best practices guide that can be useful to programmers [47]. CUDA is well suited for implement- ing the SPMD parallel design pattern [10]. Worker management in CUDA is done implicitly. That is, programmers do not manage thread creations and destructions. They just need to specify the dimension of the grid and block required to process a certain task. Workload partitioning and worker mapping in CUDA is done explicitly. Programmers have to define the workload to be run in parallel by using the function "Global Function" and specifying the dimension and size of the grid and of each block. The CUDA memory model is shown in Fig. 1. At the bottom of the figure, we see the global and constant memories. These are the memories that the host code can write to and read from. Constant memory allows read-only access by the device. Inside a block, we have the shared memory and the registers or local memory. The shared memory can be accessed by all threads in a block. The registers are independent for each thread. Finally, we would like to mention a recent initiative by Intel. This initiative is called Knights Ferry [48], [49], and is Authorized licensed use limited to: University of North Texas. Downloaded on February 19,2024 at 01:11:54 UTC from IEEE Xplore. Restrictions apply./n Bridging the Gap: Examining Parallel Programming Models for High-Performance Computing and Parallelizing Uniprocessor Simulators 1 Project Description The objective of this research is to examine the current evolution of computer architecture, specifically in relation to the increasing prevalence of processors with multiple core architectures. The work is divided into two components. Firstly, we provide a programming approach to parallelize single-processor simulators in order to efficiently mimic multi-core systems. Furthermore, we provide a comprehensive analysis of parallel programming paradigms and tools, emphasising their relevance and appropriateness for tasks associated with high-performance computing (HPC). Our study aims to enhance simulation technology and parallel computing approaches by establishing a connection between parallel programming paradigms and simula- tion environments. We aim to foster innovation in computational research and development by making a significant contribution to the advancement of parallel computing methodologies and simulation approaches. 2 Rationale In order to keep up with the increasing prevalence of multiple-core processor designs, it is crucial to de- velop simulation environments that accurately depict these architectures. Furthermore, the increasing need for high-performance computing solutions underscores the importance of understanding and using parallel programming methodologies to maximise benefits. The objective of our work is to address these significant concerns and improve parallel computing methods and simulation technologies via the investigation of par- allel programming paradigms and the parallelization of uniprocessor simulators. Our objective is to explore the complexities of parallel programming and simulation in order to meet the changing requirements of com- putational research and facilitate the development of more effective and scalable computing solutions across many fields. Our study aims to connect theoretical principles with practical applications in computational science and engineering, promoting innovation and progress. 3 What Will by implemented? The project's execution heavily relies on the development and enhancement of a programming approach to parallelize existing uniprocessor simulators. This technology enables the simulation of multiple-core architectures, providing developers and researchers with valuable tools to explore and validate innovative notions. We will conduct a comprehensive review and analysis of parallel programming paradigms and tools, evaluating their appropriateness and effectiveness for various high-performance computing applications. Through practical inquiry and assessment, our objective is to identify the primary benefits and drawbacks of different parallel programming approaches. Our goal is to contribute to the development of parallel computing frameworks that are both resilient and scalable, capable of properly utilising the computational capacity of contemporary hardware architectures. We do this via systematic experimentation and iterative improvement. The aim of our study is to provide valuable insights into the practical use and improvement of parallel algorithms. This will help in creating and using effective parallel computing solutions in many scientific and industrial fields. 1 4 Tools Needed to Implement It 5 ● Programming languages, such as C and C++, are used to implement the parallelization approach. • Employ simulation environments and frameworks such as Gem5 or Simics to examine and authenticate parallelized simulators. • Perform comprehensive research on parallel programming models and techniques by exploring databases, academic articles, and conference proceedings. • Benchmarking suites and performance analysis tools, such as Intel VTune Profiler or SPEC CPU benchmarks, are used to evaluate the effectiveness and efficiency of parallel programming models in high-performance computing (HPC) environments. • Incorporating machine learning approaches to optimise parallel algorithms and enhance performance in dynamic computing settings. • Engaging in partnerships with industrial partners and academic institutions to use state-of-the-art technology and promote multidisciplinary research in parallel computing. • Investigation of nascent hardware designs, such as GPUs and FPGAs, to use their concurrent processing capabilities and augment computational efficacy. • Creation of extensive documentation and tutorials to enable the adoption of parallel programming approaches and tools by the wider scientific community. References 1 J. Diaz, C. Muñoz-Caro and A. Niño, "A Survey of Parallel Programming Models and Tools in the Multi and Many-Core Era,” in IEEE Transactions on Parallel and Distributed Systems, vol. 23, no. 8, pp. 1369-1386, Aug. 2012, doi: 10.1109/TPDS.2011.308. keywords: Computational modeling; Parallel pro- gramming; Graphics processing unit; Message systems; Instruction sets; Multicore processing; Parallelism and concurrency; distributed programming; heterogeneous (hybrid) systems., 2 R. Kumar, V. Zyuban and D. M. Tullsen, "Interconnections in multi-core architectures: under- standing mechanisms, overheads and scaling," 32nd International Symposium on Computer Archi- tecture (ISCA'05), Madison, WI, USA, 2005, pp. 408-419, doi: 10.1109/ISCA.2005.34. keywords: Bandwidth; Computer architecture; Space technology; Delay; Power system interconnection; Joining pro- cesses;Space exploration; Computer science; Design engineering; Power engineering and energy, 3 AJ. Donald and M. Martonosi, "An Efficient, Practical Parallelization Methodology for Multicore Architecture Simulation,” in IEEE Computer Architecture Letters, vol. 5, no. 2, pp. 14-14, Feb. 2006, doi: 10.1109/L-CA.2006.14. keywords: Multicore processing; Object oriented modeling; Computational modeling; Computer architecture; Parallel programming; Computer simulation;Feedback; Parallel processing; Product development; Process planning; simulation; multicore; parallelism, 2See Answer
  • Q6:Project Final Report Submission This is the link for project final report submission. The report can be a 10 to 20 pages writeup. Please provide information of the project as follows: (1) It should have a title. (2) What is the project all about? (3) How much you have implemented already in this project? (4) What is left to be implemented that you could not do? (5) What are the difficulties and challenges you faced for this project? (6) Design and simulation figures and tables. (7) Formal references. Note: (1) The text and figures to be used should be your own. Draw the figures that you want to use. Write in your own words. (2) Use latex for your writing as it is considered better for technical writing./n Instructions Need report for the abstract which was done earlier. Task #335745 Report must be written in own words using latex and own figures to be used. No AI platforms like chatgpt/gemini. Please use own words and own figures. Need this in 10 pages | single spaced | APA formatSee Answer
  • Q7: Create a Data Center floor plan design with space to host 3 offices, rack space, heating and cooling space and show physical access and security controls. Using best practices for installing hardware racks, design one full 42u rack layout that include the following: At least five 2u servers At least one 4u storage unit At least one network switch At least one PDU At least one UPS Note any visual illustration designer like visio or lucidchart ust need simple squares, nothing fancy or 2 dimentionalSee Answer
  • Q8: 30/05/2024, 13:55 pset3-2024-spring_58_1717053712071.ipynb - Colab Spring 2024 ECE M16: PSET 3 (ver. 1.1) Student Name: LEONARD, MICHAEL Instructions (Please read !!!): 1. This notebook is personalized for you. Make sure you have downloaded the notebook that was listed against your name. 2. Deadline to submit is as listed on Gradescope. 3. You can make three submission attempts. We do not count submissions detected to have tampered with notebook cells that prevent grading (e.g. if there are more or fewer cells than expected or there is a match of cell type). For the first attempt, the Autograder will only perform some basic sanity checks on the basic format of your answer and return a score of 0. It will, however, not check or provide feedback on the correctness of your answer. For example, if the answer has to be a list of base-10 integers, then the sanity check will check whether you indeed provided 10 a list of base-10 integers. Likewise, if the answer asks for a logic expression, the sanity check will check whether your answer is indeed a valid logic expression (however, it will not check whether the variable names used are correct). For the second and third attempts, the Autograder will also report whether your answer is correct or not, but it will not tell you what the correct answer is or why your answer is wrong. 4. If we discover a bug in the Autograder, we will compensate the impacted students with an extra attempt only for the impacted question. For other questions unaffected by the bug, the score the end of third attempt will be used. 5. Please do not edit, move, or delete any of the cells. You must only write in the cells provided for your answers while adhering strictly to the format requirements. Otherwise, the Autograder cannot grade, and even though we do not count the attempt, it will cause you needless panic. 6. Most problems/subproblems will be auto-graded, and it is critically important that you strictly adhere to the formatting requirements for the answers. A few will be graded manually, and there you are allowed free-form text, either plain text or in Markdown format. 7. After every problem or subproblem, we have provided a cell where you can optionally provide a brief explanation of your solution approach. By default, we will not grade the explanation but use it if your answer is marked wrong by the autograder or the human grader, and you request a regrade. No regrade request will be entertained unless an explanation of how you derived your answer is provided. The explanation cells can accept text in Markdown format. https://colab.research.google.com/drive/1ou1VLU5PHAoq0uCKVW7TnSKGGBZOTvrN#scrollTo=Qtp5F0f2mUot&printMode=true 1/20 30/05/2024, 13:55 ✓ pset3-2024-spring_58_1717053712071.ipynb - Colab Problem 1: Sequential System Analysis [10.0 points] Consider the following sequential circuit and answer the questions in the parts that follow. When answering, use standard boolean logic operators ~, &, I, ^, and parentheses. sysclk b SO TSQ R S1 TSQ 0 R (Note: if an image doesn't show up above, please click on this link). ✓ Part 1.1: SOnext [1.5 points] Write the boolean logic expression for SO next in terms of SO and S1. Enter_your_answer_in_required_format Optional Explanation: Replace this text with your explanation. ✓ Part 1.2: S1next [1.5 points] Write the boolean logic expression for S1 next in terms of SO and S1. Enter_your_answer_in_required_format Optional Explanation: Replace this text with your explanation. Ob Z https://colab.research.google.com/drive/1ou1VLU5PHAoq0uCKVW7TnSKGGBZOTvrN#scrollTo=Qtp5F0f2mUot&printMode=true 2/20 30/05/2024, 13:55 ✓ Part 1.3: Z [1 points] pset3-2024-spring_58_1717053712071.ipynb - Colab Write the boolean logic expression for Z in terms of SO and S1. Enter_your_answer_in_required_format Optional Explanation: Replace this text with your explanation. ✓ Part 1.4: Isolated States [1.5 points] An isolated state is a state with the property that if the system starts in that state on power up, it stays there forever. Consider which of the following statements are true. 1. S0=0, S1=0 is an isolated state 2. S0=1, S1=0 is an isolated state 3. S0=1, S1=1 is an isolated state 4. S0=0, S1=1 is an isolated state 5. This system does not have an isolated state Your answer should be a list of decimal integers corresponding to the true choices. The list must not include any incorrect choices. The list may be empty (write it as []), have only one element (e.g., [1]), or may have multiple elements (e.g., [1,2,3]). Enter_your_answer_in_required_format Optional Explanation: Replace this text with your explanation. ✓ Part 1.5: Temporal Evolution of Z [4.5 points] = O we Let the system clock sysclk have a frequency of 500 MHz, and assume that at time t had a rising edge of the clock and that both the state bits SO and S1 were 0. What will be the value of Z at t = 2 ns, 4 ns, 6 ns, and 8 ns respectively? Your answer should be a comma- separated list with four elements, each of which is 0 or 1 or X (unknown), and with the first element in the list corresponding to Z at t = 2 ns, the second to to Z at t = 4 ns, and so on. E.g., https://colab.research.google.com/drive/1ou1VLU5PHAoq0uCKVW7TnSKGGBZOTvrN#scrollTo=Qtp5F0f2mUot&printMode=true 3/20 30/05/2024, 13:55 pset3-2024-spring_58_1717053712071.ipynb - Colab your answer may be [0,0,1,1] to indicate the Z=0 at t = 2 ns and 4 ns, and Z=1 at t = 6 ns and 8 ns. Note: You must get all time instant correct in order to get credit. Enter_your_answer_in_required_format Optional Explanation: Replace this text with your explanation. ✓ Problem 2: Improving System Performance [15.0 points] In all the sub-questions below, assume that the registers being used have setup time ts = 11 ps, hold time th = 9 ps, propagation delay tacq 12 = 11 ps, and no clock skew. Note thast ps means picoseconds and equals 10 s, and GOPS means billion operations per second and equals 109 operations/second. ✓ Part 2.1: Baseline Throughput [2.5 points] Consider the datapath below and assume that the combination logic has a propagation delay 319 ps. REC Combinational Logic REC N CLK CLK (Note: if an image doesn't show up above, please click on this link). What is the maximum throughput of this system, i.e. the rate at which it can process data at input X? Your answer should be a real number in units of GOPS (Giga Operations per Second). Do not include units. Your answer should be accurate to three digits after the decimal point. E.g., you answer could be 15.231. Enter_your_answer_in_required_format https://colab.research.google.com/drive/1ou1VLU5PHAoq0uCKVW7TnSKGGBZOTvrN#scrollTo=Qtp5F0f2m Uot&printMode=true 4/20 30/05/2024, 13:55 Optional Explanation: pset3-2024-spring_58_1717053712071.ipynb - Colab Replace this text with your explanation. ✓ Part 2.2: Baseline Latency [2.5 points] Continuining with the preceding, what is the input to output latency? Specifically, we are looking for the delay between when an instance of a value of ✗ was stored at a clock edge in the input register (on the left) and the value obtained after processing it gets stored at a later clock edge in the output register (on the right)? Your answer should be a number in ps. Do not include units. Enter_your_answer_in_required_format Optional Explanation: Replace this text with your explanation. ✓ Part 2.3: Throughput with Uniform Pipeline Stages [2.5 points] Now say we divide the combinational logic into three equal stages, each of which has a propagation delay of 319/3 ps. X Combinational Logic REC Combinational Logic REC Combinational Logic REC N CLK CLK CLK (Note: if an image doesn't show up above, please click on this link). CLK What is the maximum throughput of this system, i.e. the rate at which it can process data at input X? Your answer should be a real number in units of GOPS (Giga Operations per Second). Do not include units. Your answer should be accurate to three digits after the decimal point. E.g., you answer could be 15.231. REC https://colab.research.google.com/drive/1ou1VLU5PHAoq0uCKVW7TnSKGGBZOTvrN#scrollTo=Qtp5F0f2mUot&printMode=true 5/20 Enter_your_answer_in_required_format Optional Explanation:See Answer
  • Q9: Assignment 2: Consider the following instructions at the given addresses in the memory: 0x0000.1000 Load R1, 0(RO) 0x0000.1004 Add R2, R1, #5 0x0000.1008 And R3, R2, #0x0000.0010 0x0000.100C Store R3, 4(R4) 0x0000.1010 Add R5, R2, R3 Initially, RO and R4 contain 0x0000.2000 and Ox1000.2000, respectively. These instructions are executed in a computer that has a five-stage pipeline as described in Chapter 6. The first instruction is fetched in clock cycle 1. (a) Assume all memory access operations are cache hit. Complete the diagram that represents the flow of the instructions through the pipeline. How many cycles does it take to execute all the instructions if data forwarding is allowed? Also add marks to the diagram to indicate where data forwarding is used. Cycle 1 2 3 Load R1, 0(RO) F D... Add R2, R1, #5 F And R3, R2, #0x0000.0010 Store R3, 4(R4) Add R5, R2, R3 (b) Assume only memory accesses to addresses 0x1000.0000-0x3000.0000 are cache hit. Accesses to all other addresses are cache miss, which take 3 cycles to complete. Repeat Problem (a) to draw a diagram that represents the flow of the instructions through the pipeline. How many cycles does it take to execute all the instructions if data forwarding is allowed? Also add marks to the diagram to indicate where data forwarding is used.See Answer
  • Q10: Assignment 1: Consider the following instructions given below. I1: Add R1, R2, R3 12: Add R2, R1, R4 13: Add R1, R1, R4 These instructions are executed in a computer that has a five-stage pipeline as described in Chapter 6. The first instruction is fetched is clock cycle 1. The clock period for the pipelined processor is specified below for the processor with and without forwarding logic. Forwarding With Forwarding Clock Period 100 psec Without Forwarding 150 psec (a) Identify the data dependencies and their type (i.e. RAW, WAR, and WAW). For example, use the format: 14 →> 12, Register RO, RAW. To identify a RAW data dependency between instructions 12 and 14 due to register RO. (b) Assume that there is NO forwarding logic in the processor. Identify the RAW data hazards. For example, use the format: 13 →12, Register R4. To identify a RAW data hazard between instructions 12 and 13 due to register R4. (c) Assume that there is forwarding logic in the processor. Identify the RAW data hazards. Use the same format described above. (d) What is the total execution time (ET) of this code segment on the processor without forwarding logic and with forwarding logic? What is the speedup achieved by adding forwarding logic to a pipeline that did not have it? i. Without forwarding logic: You can use the flowing diagram to calculate the total clock cycles : Cycle Add R1, R2, R3 1 2 3 F D Add R1, R2, R3 Or: use the flowing table: IF ID EX MEM WB 1 I1 2 3 4 5 6 7 8 9 10 11 ii. With forwarding logic:See Answer
  • Q11:3. Design a Static RAM (SRAM) memory unit of size 3 words, each word is of size 4 bits. Show the details of all digital components used in the design including the memory cells. (30 points)See Answer
  • Q12:2. Design a common bus for a basic computer that has 3 registers of size 5 bits each. The solution should use three-state buffers. (30 points)See Answer
  • Q13:1. A digital computer has a common bus system for 16 registers of size 32 bits each. The bus is constructed with multiplexers. (10 points) a) How many selection inputs are there in each multiplexer? b) What size of multiplexers are needed? c) How many multiplexers are there in the bus?See Answer
  • Q14:3. Step: 3a - The heart of the mini-project. VM24 Only - Due Sunday, 4/7 by 11:59pm Design the microarchitecture interpreter for executing/interpreting the above macro-program by: a. First, design the virtual hardware (Level 1) - CPU - for executing each micro-architecture instruction based on the Instruction Set given to you below. The design of your virtual machine (CPU) would be a pseudocode in C, C++, or Java programming language. Your virtual machine should include all the necessary registers and other data structures needed for the task. (See the instruction set and instruction format in section 4, and the types of special and general-purpose registers referenced in the above assembly code.) b. The OS machine level (sitting on top of your ISA-level) would be a driver', a pseudocode of the master program, that simply allocates space (RAM memory) for the various parts of the program listed in section 2 above. (The OSML doesn't accomplish much, at this point - no fancy memory management or scheduling, or synchronization - since you have only one program/job's execution to simulate.) c. All the variables, registers, etc. in the driver and the virtual CPU define the state of the VM24 computer. For example, the PC is part of the state. Therefore, it is instructive to define a data structure, called PCB (process control block) to maintain the state of the computer. The following schematic, for the VM24 architecture, is a simple guide for you to design and implement each component of the architecture. Note: You have a choice between C, C++, or Java for the pseudocode. The focus is on the OS driver, the CPU and its components, the RAM (skip the 'Disk') and have the Loader load the 'microprogram' into the RAM. [Note: the 32-bit microinstructions will assist in the instruction 'Decode' stage.] Step 2b above should guide you in getting the Decoder's logic right. The Execute will serve as the ALU (made of functions', or opcodes) which will 'interpret' each instruction of the program listed in the Assembly/Hex form. The Long-Term and the Short-Term/Dispatch components are not needed, except the PCB which carries the 'state' of the VM24. There will be no "Context-Switching in this exercise. C Pr/nDecode (info) (0) Fech CPU PC+1 PC=0 PC РСВ 1d 8-23 State 8 19 080 OS driver Scheduler 10 Execute (para) Opcode...... [memory (info)] Program File 00000 Long scheduler disk Loader Effective_addr) //JOB 1EB 2C15212 25 words ABCDEFA person Instruct data 02C12345 Data BCC 2 DATA portion END Job #2 JOB 2 CF 19 (: 3 2048 job1 job 2 0-18 (memory) RAM 2 1 1024 8 PCBi Short Scheduler dispatcher Context Switch RQ 2 1 4-byte word (8 hex-chars) copy 16 PCBI CPUI г2 General Purpose Cpu reg PCREG register Г1 PC2See Answer
  • Q15:INSTRUCTIONS Create an abstract-Word Limit Is 650-700 Please provide information of the project as follows: It should have a title. What is the project all about? Why do you want to do this specific as a project? What will you implement? What tools needed to implement it? Formal References. Note : Can Work On Any IEEE Related PaperSee Answer
  • Q16:3- A MOV assembly instruction copies the value of the source register to the destination register. What is the value of the destination register r1 after the following instruction is executed? Memory Address 0x08000166 *** Assembly Instruction MOV rl, pc ***See Answer
  • Q17:CS370 Lab Assignment #2 Goal: The purpose of this project is to demonstrate the use of Boolean logic and build basic circuits. Note that this lab assignment is group assignment with each group having either 2 students or 3 students. Problem Statement: The predominant storage inside computer systems are on disks drives. SCSI disks are the standard disks in most Unix Workstations from Sun, HP, SGI, and other vendors. They are also the standard disks in Macintoshes and Higher-end Intel PC's, especially network servers. Consider the Wide Ultra4 SCSI which transfers data packets in 16-bit bursts at 160 MHz with a maximum throughput of 320 MB/sec. The data transfers at higher rates can result in random-noise pulse changes from a 0 to 1 and 1 to a 0. As the speed of processors and electronic communications increases, these parity flips become more prevalent and the inability to detect when these errors occur can be fatal. As a Design Engineer you have been requested to create system for the transmission of these 8-bit packets from an I/O Controller to Memory using Error Correcting Code over 12-bit data bus line. Wide SCSI contains a 68bit bus; however for the sake of simplification we are only concerned with the data bits. The other bits in the SCSI bus are for bus arbitration, synchronization, power management, etc. In this project, we will use even parity. Hex Displays Memory I/O Controller Transmission Vectored Bit: A 4-Bit Parity Vector (P₁-P4) are interlaced with the 8-bit Data Vector (D₁ D₂): P₁ P₂ D₁ P3 D₂ D3 D4 P4 Ds Dg D7 Ds 1) Create an ECC Generator, at the I/O Controller from the 8-bit Data Vector. The output of the ECC Generator will be the 4-Bit Parity Vector. 2) Construct a 12-bit Data Transmission bus to send the binary data and parity bits over to Memory. 3) Construct an ECC Detector at Main Memory that corrects for single bit errors. Generally, an interrupt/error handler is used to handle errors from the OS, for this exercise we will use 3 Hex displays, 2 for data and 1 for an error status, for diagnostic purposes: In the event that no error has occurred, your design must display the data transferred using the 2 Hex data displays and a "0" as an error status. . For single bit transmission errors, your system must correct the error and display the data along with "C" in the 3rd Hex Display. For multiple bit transmission errors, your design must display "E" in the error status display.See Answer
  • Q18:You have the following circuit: A C- محمد D- 1) Code it using behavioral modeling. Show a screenshot of your code. -Y 2) Write a testbench for the code of Part 1. Show screenshots of the code and the waveform, clearly showing the results 3) Code it using structural modeling. Show a screenshot of your code. 4) Write a testbench for the code of Part 3. Show screenshots of the code and the waveform, clearly showing the result. use EDA playground web to do this assignment | https://www.edaplayground.com/See Answer
  • Q19:Description of the IS Architecture, Software and Database Components, and Hardware Architecture The information system designed for this case should connect docking stations, as well as manage user registrations and billing. We learned about different architecture models (e.g. centralized, distributed, and cloud). Directions In a Word document, provide the following information: Propose the IS architecture for the IS for the bike sharing system. Explain the rationale behind the choice List all the software, and database, and hardware components; Draw a architecture of the system showing the components using Word shapes, use arrows to show business process and major data flow between them. APA Format, 550 words, Double spacing. References will not be count in 550 words.See Answer
  • Q20:Summer.2023 me labus nouncements dules signments ades scussions om Pro 1.3 40 Use the following module interface (which you can find in the provided file): module NextPCLogic(NextPC, CurrentPC, SignExtImm64, Branch, ALUZero, Uncondbranch); input [63:0] CurrentPC, SignExtImm64; input Branch, ALUZero, Uncondbranch; output [63:0] NextPC; reg [63:0] tmp; /* write your code here */ endmodule Branch (CBZ) is true if the current instruction is a conditional branch instruction, Uncondbranch is true if the current instruction is an Unconditional Branch (B), and ALUZero is the Zero output of the ALU. A template and testbench are provided (If you work on Vivado, comment out the include command at the very top). Complete the next pc logic; add a few test cases to the testbench to improve it. Demonstrate your program to the TA. Attach a zip/tar file containing your completed module, testbench and a screenshot of the waveform. NextPClogic.v↓↓ NextPClogicTest.v Upload Choose a File/ny mer.2023 IS ncements es ments S Electrical and Co... sions Pro 1.3 40 Home | Howdy Question 2 Outlook PC Add WebAssign-LOG... C Get Homework He... Write a behavior model to calculate the next PC for an instruction. It will use information from the processor control module and the ALU to determine the destination for the next PC. It will contain two adders for calculating the two possible NextPC and choose between the two possibilities using the logic depicted below. memory Regi:00 studion Dif [Parucion H - Uncondbranch Branch ""!! data Reg Netflix ▸YouTube M Gmail. Shift left 2 Add ALU result 20 pts Zero 2 MapsSee Answer

TutorBin Testimonials

I found TutorBin Computer Organisation And Architecture homework help when I was struggling with complex concepts. Experts provided step-wise explanations and examples to help me understand concepts clearly.

Rick Jordon

5

TutorBin experts resolve your doubts without making you wait for long. Their experts are responsive & available 24/7 whenever you need Computer Organisation And Architecture subject guidance.

Andrea Jacobs

5

I trust TutorBin for assisting me in completing Computer Organisation And Architecture assignments with quality and 100% accuracy. Experts are polite, listen to my problems, and have extensive experience in their domain.

Lilian King

5

I got my Computer Organisation And Architecture homework done on time. My assignment is proofread and edited by professionals. Got zero plagiarism as experts developed my assignment from scratch. Feel relieved and super excited.

Joey Dip

5

TutorBin helping students around the globe

TutorBin believes that distance should never be a barrier to learning. Over 500000+ orders and 100000+ happy customers explain TutorBin has become the name that keeps learning fun in the UK, USA, Canada, Australia, Singapore, and UAE.