Showing posts with label checkpoints. Show all posts
Showing posts with label checkpoints. Show all posts

Monday, March 12, 2012

Fast Forwarding and Simpoint support in MARSS

One of the most requested feature in MARSS has been fast-forwarding N instructions before starting simulation. Recently we have developed some hooks into QEMU's translation logic to count number of instructions emulated. With this logic we are now able to support fast-forwarding and simpoint in MARSS. Both of these features are available in features branch right now.

Fast-Forwarding
Following new simconfig options are added for fast-forwarding support in MARSS.
  • -fast-fwd-insns N: Fast-forward N number of total instructions. This includes user level and kernel level instruction across all emulated CPUs. After specified limit is reached it will switch to simulation mode.
  • -fast-fwd-user-insns N: Fast-forward N number of user level instructions. This mode will emulate kernel level instructions but doesn't count them. After specified user level instructions are executed it will switch to simulation mode and it will simulate both user and kernel level instructions.
  • -fast-fwd-checkpoint CHK_NAME: Create a checkpoint named 'CHK_NAME' after fast-forwarding specified amount of instructions. It will kill simulation instance after a checkpoint is created.
As mentioned above the default behavior is to switch to simulation mode after specified instructions are emulated. Because of non-deterministic behavior of kernel level execution in emulation mode we also support creating a checkpoint after fast-forwarding so users can use fast-forwarding only once and then rely on checkpoint to start simulation at specific RIP everytime.

For multicore emulation/simulation, fast-forwarding specified amount of instruction is little complicated because many times only few of CPUs are executing any code and others are in idle mode. At first MARSS equally divides total number of instructions to emulate and allocate them to all CPUs in machine. When part of those CPUs are in idle mode then MARSS will re-allocate portion of remaining instruction count from idle CPUs to non-idle CPUs to reach to specified instructions limit.

Simpoint
Simpoints has been one of the most used and reliable method in computer-architecture research to simulate part of applications that is representative of full application run. We decided to implement support for Simpoints to evaluate performance of different applications with real-hardware, but more on that later. MARSS uses simpoint file to create checkpoints after specified instructions are emulated. To create checkpoints based on simpoints and how to use weights file and mstats.py to calculate weighted IPC please refer to wiki page on Simpoints.

Counting Emulated Instructions
QEMU's emulation engine - TCG is designed to be fast which first convert instruction into list of micro-instructions and then convert these micro-instructions into binary buffer that can be executed and re-executed without any change. This makes little bit difficult to directly count the number of instructions emulated.

Recently we developed support for Simpoints in which we need to count number of instructions emulated and create a checkpoint after specific number of instructions. During this development we figured out a way to add hooks into each translated binary block to count number of instructions emulated at run-time. We added a counter to each CPU context which is set to number of instructions a CPU context is allowed to execute. In each translated block we added a hook that decrements CPU context's instruction counter by number of instructions in the block. When the counter reaches to 0 we switch to simulation mode or create a checkpoint based on user's configuration options.

Tuesday, September 27, 2011

Creating a Barrier between Processes with SysV Semaphore

Running deterministic simulations gets too much complicated when you are dealing with more than one processes. In each simulation run you need to make sure that processes are assigned to same CPUs all the time and kernel scheduler does not introduce any randomness. For single process with multiple threads its easy to synchronize them using thread level synchronizations like mutex, barriers etc.. But with multiple processes things get little complicated without proper allocation of resources to each process and its threads. Also the point when we create a checkpoint is important because we need to make sure that all resource allocation by kernel is done and benchmarks are ready to start their Region-Of-Interest (ROI).

In this post I'll explain how to use System V IPC mechanism to create a barrier between multiple processes. The general idea is to wait on a barrier before each thread starts executing ROI and once all threads reach this point we create a checkpoint and run our simulations from this checkpoint only. The trick is to use SysV semaphores as a barrier across processes. It provides a semaphore operation mode where each process/thread wait for semaphore value to become 0. In short we initialize a semaphore to total number of threads and each thread will decrement the semaphore and wait till its value reaches to 0.

Create/Get a Semaphore
First we create a semaphore using semget. As described in man page, if we pass the key parameter to a specific value it will create a process shared semaphore which can be accessed by other processes using semget.
static int SEM_ID = 786;

int get_semaphore ()
{
    int sem_id;

    sem_id = semget(SEM_ID, 1, IPC_CREAT | 0666);

    if (sem_id == -1) {
        perror("get_semaphore: semget");
        exit(1);
    }

    return sem_id;
}
As highlighted in line 7 we use a common SEM_ID to get the semaphore. get_semaphore function returns the semaphore id which will be used in later function to set semaphore value and perform operations on it.

Set Semaphore's Value
Once we get the semaphore id then we call semctl to set the value of the semaphore. Remember that we must call this function only once.
int set_semaphore (int sem_id, int val)
{
    return semctl(sem_id, 0, SETVAL, val);
}


Decrement Semaphore
Now when a thread/process reaches the start of ROI region they should decrement the semaphore value and wait till its value becomes zero. To decrement the semaphore we use semop function. semop uses a pointer to a sembuf structure which specifies the operation to perform on semaphore.
void decrement_semaphore (int sem_id)
{
    struct sembuf sem_op;

    sem_op.sem_num  = 0;
    sem_op.sem_op   = -1; /* <-- Specify decrement operation */
    sem_op.sem_flg = 0;

    semop(sem_id, &sem_op, 1);
}
As highlighted in line 6 we set sem_op variable to -1 which will tell the kernel to perform atomic decrement operation on the semaphore.

Wait for other Threads/Processes
Now we just have to wait till all the processes reach to our barrier. For that we again use semop function. This time we set sem_op variable's value to 0 to tell the kernel that wake-up this thread/process once semaphore value is 0.
void wait_semaphore (int sem_id)
{
    struct sembuf sem_op;

    sem_op.sem_num  = 0;
    sem_op.sem_op   = 0;
    sem_op.sem_flg = 0;

    semop(sem_id, &sem_op, 1);
}
This function call will block until all the threads have decremented the semaphore and its value is 0.

I have put all these function into a single file: sem_helper.h.

Putting it all Together
Now we create a small application (set_semaphore.cpp) that will create and initialize the semaphore to specific value.
#include 
#include 
#include 
#include 

#include "sem_helper.h"
#include "ptlcalls.h"

using namespace std;

int main (int argc, char** argv)
{
    int sem_id;
    int sem_val;
    int rc;

    if (argc < 2) {
        cout << "Please specify the initial semaphore value.\n";
        return 1;
    }

    /* Get semaphore */
    sem_id = get_semaphore();

    /* Retrive semaphore value from command line arg */
    sem_val = atoi(argv[1]);

    /* Set semaphore value */
    rc = set_semaphore(sem_id, sem_val);
    if (rc == -1) {
        perror("set_semaphore: semctl");
        return 1;
    }

    /* Now wait for semaphore to reach to 0 */
    wait_semaphore (sem_id);

    /* All threads will be in ROI so either create a checkpoint or
     * switch to simulation mode. */
    ptlcall_checkpoint_and_shutdown("checkpoint_1");

    return 0;
}
Compile this code into an application and run it as shown below. (We run it in background mode because once all threads reach to barrier we call the 'ptlcall' to create a checkpoint.)
$ ./set_semaphore NUM_THREADS &
In next step we modify our benchmarks to use this semaphore functions to wait at the barrier. Just before ROI begins write the following code.
#include "sem_helper.h"

void sync_all_processes ()
{
    int sem_id;
    int rc;

    /* Get semaphore */
    sem_id = get_semaphore();

    /* Decrement semaphore value */
    decrement_semaphore (sem_id);

    /* Now wait till semaphore value reaches 0 */
    wait_semaphore (sem_id);
}
Tip: For Parsec Benchmarks you can modify the hook to call the sync_all_processes.