Launching the Open-Source Rocket Chip Generator!

We are very excited to announce the alpha release of our Rocket chip generator. This generator toolkit can be used to create instances of our high-performance, energy-efficient Rocket processor suitable for both high-speed simulation and full synthesis. We have provided a collection of components that go well beyond simple pipeline RTL in order to allow you to generate a complete Rocket implementation, including a memory subsystem.

What is Rocket?

Rocket is a 5-stage single-issue in-order pipeline that executes the 64-bit scalar RISC-V ISA (see the pipeline diagram below). The scalar datapath is fully bypassed but carefully designed to minimize the impact of long clock-to-output delays of compiler-generated SRAMs in the caches. For example, we have moved the branch resolution to the M stage to reduce the critical D$ bypass path. A branch target buffer, a two-level branch predictor, and a return address stack together mitigate the increased performance impact of these control hazards.  Rocket implements an MMU that supports page-based virtual memory and is able to boot modern operating systems, including Linux.  Both caches are virtually indexed and physically tagged with parallel TLB lookups.  The data cache is non-blocking, allowing the core to exploit memory-level parallelism.

Rocket has an optional IEEE 754-2008-compliant FPU, which can execute single- and double-precision floating-point operations, including fused multiply-add (FMA), with hardware support for subnormals and other exceptional values.  The fully-pipelined double-precision FMA unit has a latency of three clock cycles.

rocket-pipeline-diagram

Note: The description above is an excerpt from our ESSCIRC paper “A 45nm 1.3GHz 16.7 Double-Precision GFLOPS/W RISC-V Processor with Vector Accelerators“.

What Does the Generator Do?

The Rocket chip generator produces parameterized RTL for the Rocket core and associated caches, glues it to an uncore/memory system, and then transforms all of the above into one of the following:

  1. High-speed C++ simulator: The generated C++ simulator is faster than a Verilog simulator, and is also able to generate vcd waveform dumps for debugging.
  2. Verilog optimized for FPGA: We also released scripts to map the generated Verilog down to a ZYNQ FPGA. Please take a look at the open-sourced fpga-zynq repository.
  3. Verilog optimized for VLSI: We have taped out three 28nm test chips and five 45nm test chips with the Rocket chip generator. Unfortunately, we cannot open source VLSI scripts due to NDA restrictions.

Who Should Use the Generator?

People who want a RISC-V core for VLSI or FPGA implementation. If you just want it for that purpose, as is, then little work is required of you. Please note that Rocket is designed and optimized for VLSI implementations, and therefore will not result in the best FPGA resource utilization relative to a dedicated soft core design. However, it is a real processor that runs our entire software stack (including Linux and user-level applications such as cross-compiled Python).

With a little effort, you can understand the provided parameters and change them to better utilize your available HW resources. With more effort, you can learn Chisel and make changes to the design. In the future, we’ll make it easier to integrate custom accelerators.

What Are the Next Steps?

Download the Rocket chip generator from github at this link: https://github.com/ucb-bar/rocket-chip. The README addresses the specific contents of the rocket-chip repository, how to generate particular Rocket implementations, and how to parameterize a Rocket core.

About our Dhrystone benchmarking methodology

This is an extended version of Krste’s comment on the RISC-V EE Times article about our Dhrystone benchmarking methodology.

We have reported a Dhrystone score of 1.72 DMIPS/MHz for the Rocket core here. We pulled the Dhrystone comparison together quickly, as we kept getting asked about how we compared to ARM cores and these were the only publicly available numbers we could easily compare against. We didn’t spent a lot of time on it, as we’re not particularly interested in “Dhrystone Drag Racing” with minimal stripped-down cores. Basically, we sized the caches to match ARM’s configuration publicly available on their website and removed our vector floating-point unit to make a fairer comparison with the A5, which was also configured without an FPU or vector unit. We didn’t strip out a lot of other stuff that we could have.

Here’s a more detailed specification of the Rocket core we’ve used:

  • The Rocket core implements RV64IMA, i.e., base integer, integer multiply/divide, and atomic operations (which are quite extensive in RISC-V and go unused in Dhrystone). Our registers are twice as wide (64 vs 32) and we have twice as many user registers (32 versus 16) as ARM. The 64-bit width does help Dhrystone, but also lots of other code, and they are obviously included in our area number.
  • The instruction cache was 16KB 2-way set-associative with 64-byte lines, blocking on misses.
  • The data cache was also 16KB 2-way set-associative with 64-byte lines, but because it was designed to work with our high-performance vector unit, it’s non-blocking with 2 MSHRs, 16 replay-queue entries, and 17 store-data queue entries. None of these help Dhrystone, which never misses in the caches.
  • When we compared numbers with and without caches, we weren’t sure what ARM left out, so we only removed the SRAM tag and data arrays and left in all of the above cache control logic in our core area. The MMU has 8 ITLB and 8 DTLB entries, fully associative, and the MMU has a hardware page-table walker. Obviously, the hardware page table walker doesn’t help Dhrystone.
  • The branch prediction hardware is a BTB with 64 entries, a BHT with 128 entries, and a RAS with 2 entries. This amount of branch prediction helps Dhyrstone, but helps a lot of other code, too.

We made sure to adhere to the guidelines in ARM’s Dhrystone benchmarking methodology when compiling the code.  More specifically, we use the following lines to invoke the compiler:

$ riscv-gcc -c -O2 -fno-inline dhry_1.c
$ riscv-gcc -c -O2 -fno-inline dhry_2.c
$ riscv-gcc -o dhrystone dhry_1.o dhry_2.o

The disassembled benchmark code can be found here.

Our standard C library does include hand-optimized assembly, and does make use of all 64-bits (of course!), but we also did this for functions not used by Dhrystone, as an optimized standard library helps all code.

Like ARM, we didn’t actually fabricate this version of the RTL, but we have fabricated and measured enough variants in different processes to be confident in our layout results.

Overall, we’re pretty sure it’s a reasonable comparison, though we’re not completely sure about all the details in ARM’s result to make sure we’re being fair. But you don’t need to take our word for it: you can checkout our Rocket core generator and replicate the same experiment on your end!

RISC-V at ESSCIRC-2014

We have presented our paper “A 45nm 1.3GHz 16.7 Double-Precision GFLOPS/W RISC-V Processor with Vector Accelerators” at the 40th European Solid-State Circuits Conference, which was held at Venice, Italy. This paper details our 45nm test chip, which has two 64-bit RISC-V Rocket scalar cores, each with a Hwacha vector accelerator attached to it. The paper and the talk will be available online shortly.

RISC-V Analysis by Adapteva Founder

Andreas Olofsson is the founder of Adapteva and the creator of the Epiphany architecture and Parallella open-source computing project.  He has published his own analysis of the RISC-V ISA. Quoting the conclusion:

“The RISC-V architecture is not revolutionary, but it is an excellent general purpose architecture with solid design decisions. The true breakthrough here is really the open source licensing model and the maturity of the design as compared to most other open source hardware projects. I am personally very enthusiastic about the kinds of low cost systems that will be built around this RISC-V architecture going forward. A royalty free 64-bit RISC-V core would have a raw silicon cost of a couple of cents in current CMOS process nodes. Now that is exciting!”

 

ARM’s rebuttal to RISC-V: “The Case for Licensed Instruction Sets”

MICROPROCESSOR Report has publicly released an extended version of our technical report “Instruction Sets Should Be Free: The Case For RISC-V” and ARM’s corresponding rebuttal.  The RISC-V publication webpage has been updated with the following links:

RISC-V at HotChips-26

The RISC-V team was out in force at the HotChips-26 conference manning a sponsor booth.

photo1

Yunsup arrives early on Sunday to set up the RISC-V booth.

photo2

Alongside a bunch of cool giveaways, including RISC-V buttons and bumper stickers, we had several demo boards at the conference. From left to right: a dual-core Rocket+Hwacha system in IBM’s 45nm SOI process, running up to 1.35GHz, a single core Rocket+Hwacha system in ST 28nm FDSOI process running down to 0.45V, and an FPGA prototype running on a Xilinx Zybo board.

photo3

The Berkeley RISC-V team pose for a group shot at the end of the conference. From left to right: Steven Bailey, Henry Cook, Sagar Karandikar, Palmer Dabbelt, Krste Asanovic, Adam Izraelevitz, Colin Schmidt, Yunsup Lee, Andrew Waterman, Brian Zimmer, Scott Beamer, David Patterson.

RISC-V just got a new logo!

As we were getting ready for HotChips, we realized we were missing something very important: A logo!

Thanks to the creative designers at 99designs, we were able to get a pretty good logo in a week. The symbol visualizes RV, which we often use to abbreviate RISC-V when naming an ISA variant. The logo comes in two layouts. First, here’s the tall RISC-V logo:

riscv-symbol-text-standard-tall-square

The wide RISC-V logo also looks great:

riscv-symbol-text-standard-wide

 

For dark backgrounds, we substitute the blue color with white.

riscv-symbol-text-by-wide

riscv-symbol-text-bw-wide