Friday, October 19, 2007

Saved By Chappelle

The last three weeks have been a hell. If it could go wrong it went wrong. It started with the death of my workstation followed quickly by a fender bender. But I'm not going to dwell too much on the fender bender because I got off easy. The short version is, the moron that ran into me didn't get beaten to a pulp, thus I'm not writing this from a prison cell (thank you Dave Chappelle). My car [miraculously] only sustained minor scratches to the rear bumper (you should have seen the other guy), which I'm not going to fix it because it means repainting the bumper, which would make my car visually lopsided. That happened with my very first car [89 Honda CRX Civic]. The insurance company wouldn't pay for the whole thing to be painted and I didn't have the money to do it out of pocket. So the passenger side door and front panel was a different shade of yellow than the rest of the car. It drove me crazy. My current car is a '06 Ford Mustang GT, black. I [still] plan on tricking it out so the rear bumper was going to go anyway. So it's pointless to paint the bumper now and suffer unnecessarily.

What really pissed me off about the accident was, (a) it was completely avoidable (I have no idea why people think a yellow light means "speed up") and (b) the car just celebrated it's first birthday. The thing that kept me from losing it was the, "When Keeping it Real Goes Wrong", skits from the Chappelle Show. First there was contact; then there was blinding rage; then out of nowhere, Dave Chappelle. Weird! Network TV should run them as public service announcements. I think it would benefit Type A personalities.

Tuesday, September 18, 2007

[BUG] JE never stops logging ...

This is entry is a partial repost of a message I posted to Oracle's Berkeley DB JE Forum. The forum software does not allow for the proper formatting of source code and I personally hate reading unformatted source code. Therefore, I have reposted it here so people like me can't read the code right off the page.

The code

import java.io.File;
import java.lang.reflect.Field;
import java.util.List;
import java.util.Random;

import com.sleepycat.je.Cursor;
import com.sleepycat.je.Database;
import com.sleepycat.je.DatabaseConfig;
import com.sleepycat.je.DatabaseEntry;
import com.sleepycat.je.DatabaseException;
import com.sleepycat.je.Environment;
import com.sleepycat.je.EnvironmentConfig;
import com.sleepycat.je.OperationStatus;
import com.sleepycat.je.Transaction;
import com.sleepycat.je.evictor.Evictor;

import it.unimi.dsi.fastutil.ints.IntOpenHashSet;

public class CrushPunyJE
{
    private static final int NUM_DOMAINS = 1000;

    private static final int RECIPIENTS_PER_DOMAIN = 1000;

    public static void main(String[] args) throws Exception
    {
        crush(new File(args[0]));
    }

    private static void crush(File envHome) throws Exception
    {
        MyDbI dbi = MyDbI.openToFill(envHome);
        Cursor queue = dbi.queue().openCursor(null, null);
        Cursor recipients = dbi.recipients.openCursor(null, null);
        Cursor unsent = dbi.unsent.openCursor(null, null);
        DatabaseEntry ename = new DatabaseEntry();
        DatabaseEntry dname = new DatabaseEntry();

        try
        {
            Random rand = new Random();
            byte[] buff = new byte[512];
            for (int a = 0; a < NUM_DOMAINS; a++)
            {
                rand.nextBytes(buff);
                dname.setData(buff);
                for (int z = 0; z < RECIPIENTS_PER_DOMAIN; z++)
                {
                    rand.nextBytes(buff);
                    ename.setData(buff);
                    if (OperationStatus.SUCCESS == unsent.putNoDupData(dname, ename))
                    {
                        recipients.put(dname, ename);
                        queue.put(dname, ename);
                    }
                    else
                        z--;
                }
            }
        }
        finally
        {
            for (Cursor c : new Cursor[]{recipients, unsent, queue})
            {
                try
                {
                    c.close();
                }
                catch (DatabaseException e)
                {
                    e.printStackTrace();
                }
            }
            dbi.close();
        }

        final MyDbI dbi2 = MyDbI.open(envHome);
        final Runnable runnable = new Runnable()
        {
            public void run()
            {
                Evictor evil = (Evictor)getField(getField(dbi2.env(), "environmentImpl"), "evictor");
                while (true)
                {
                    evil.runOrPause(true);
                    try
                    {
                        Thread.sleep(10);
                    }
                    catch (InterruptedException e)
                    {
                    }
                }
            }
        };
        Thread t = new Thread(runnable);
        t.setDaemon(true);
        t.setPriority(Thread.MAX_PRIORITY);
        t.start();
        dbi2.clearQueue().sync();
        dbi2.close();
    }

    private static Object getField(Object o, String fieldName)
    {
        Field field = getField(o.getClass(), fieldName);
        if (null == field)
            throw new AssertionError("Field '" + fieldName + "' not found.");
        else
        {
            try
            {
                return field.get(o);
            }
            catch (IllegalAccessException ex)
            {
                throw new AssertionError("IllegalAccessException thrown while accessing: ".concat(fieldName));
            }
        }
    }

    private static Field getField(Class co, String name)
    {
        for (Class stop = Object.class; co != stop; co = co.getSuperclass())
        {
            for (Field field : co.getDeclaredFields())
            {
                if (name.equals(field.getName()))
                {
                    field.setAccessible(true);
                    return field;
                }
            }
        }
        return null;
    }
    //====================================================================================================================//
    //====================================== Inner Class Definitions Start Here ==========================================//
    //====================================================================================================================//

    private static class MyDbI
    {
        /**
         * A small cache of prime numbers.
         */
        private static final IntOpenHashSet PRIMES = new IntOpenHashSet();

        /**
         * The maximum number of lock tables.
         */
        private static final int MAX_LOCK_TABLES = 523;

        /**
         * The databases.
         */
        public final Database recipients, unsent;

        /**
         * The delivery queue.
         */
        private static final String QUEUE_DB = "queue";

        /**
         * The database name of the unsent recipients.
         */
        private static final String UNSENT_DB = "unsent";

        /**
         * The database name for the recipient list database.
         */
        private static final String RECIP_DB = "recipients";

        private static final int CACHE_SIZE = 1024 << 10 << 4;

        public Database queue;

        /**
         * The database environment object.
         */
        private final Environment env;

        private final boolean dw, readonly;

        private MyDbI(File home, EnvironmentConfig ecfg, DatabaseConfig dbc) throws DatabaseException
        {
            dw = dbc.getDeferredWrite();
            readonly = dbc.getReadOnly();
            dbc.setSortedDuplicates(true);

            env = new Environment(home, ecfg);

            Transaction txn = ecfg.getTransactional() ? env.beginTransaction(null, null) : null;

            if (dbc.getAllowCreate())
            {
                queue = env.openDatabase(txn, QUEUE_DB, dbc);
                unsent = env.openDatabase(txn, UNSENT_DB, dbc);
                recipients = env.openDatabase(txn, RECIP_DB, dbc);
            }
            else
            {
                List names = env.getDatabaseNames();
                queue = names.contains(QUEUE_DB) ? env.openDatabase(txn, QUEUE_DB, dbc) : null;
                unsent = names.contains(UNSENT_DB) ? env.openDatabase(txn, UNSENT_DB, dbc) : null;
                recipients = names.contains(RECIP_DB) ? env.openDatabase(txn, RECIP_DB, dbc) : null;
            }

            if (null != txn)
                txn.commit();
        }

        /**
         * @return The Environment object that backs this DbI.
         */
        public Environment env()
        {
            return env;
        }

        /**
         * Get the queue database.
         *
         * @return The queue database.
         */
        public synchronized Database queue()
        {
            return queue;
        }

        public synchronized MyDbI clearQueue() throws DatabaseException
        {
            DatabaseConfig config = queue.getConfig();
            queue.close();
            env.truncateDatabase(null, QUEUE_DB, false);
            queue = env.openDatabase(null, QUEUE_DB, config);
            return this;
        }

        /**
         * Synchronizes {@link #unsent} with {@link #queue}.
         *
         * @return this
         *
         * @throws DatabaseException If there is a database error.
         */
        public synchronized MyDbI sync() throws DatabaseException
        {
            if (readonly)
                throw new IllegalStateException("read-only mode");

            DatabaseEntry k = new DatabaseEntry();
            DatabaseEntry v = new DatabaseEntry();
            Cursor uc = unsent.openCursor(null, null);
            try
            {
                if (dw)
                {
                    while (OperationStatus.SUCCESS == uc.getNext(k, v, null))
                        if (OperationStatus.SUCCESS != queue.putNoDupData(null, k, v))
                            assert false : "Duplicate email addresses not allowed in queue.";
                }
                else
                {
                    boolean commit = false;
                    Transaction txn = env.beginTransaction(null, null);
                    try
                    {
                        Cursor qc = queue.openCursor(txn, null);
                        try
                        {
                            while (OperationStatus.SUCCESS == uc.getNext(k, v, null))
                                if (OperationStatus.SUCCESS != qc.putNoDupData(k, v))
                                    assert false : "Duplicate email addresses not allowed in queue.";
                            commit = true;
                        }
                        finally
                        {
                            qc.close();
                        }
                    }
                    finally
                    {
                        if (commit)
                            txn.commit();
                        else
                            txn.abort();
                    }
                }
            }
            finally
            {
                uc.close();
            }
            return this;
        }

        /**
         * Synchronizes {@link #unsent} with {@link #queue}.
         *
         * @return this
         *
         * @throws DatabaseException If there is a database error.
         */
        public synchronized MyDbI sync2() throws DatabaseException
        {
            if (readonly)
                throw new IllegalStateException("read-only mode");

            DatabaseEntry k = new DatabaseEntry();
            DatabaseEntry v = new DatabaseEntry();
            Cursor uc = unsent.openCursor(null, null);
            try
            {
                if (dw)
                {
                    while (OperationStatus.SUCCESS == uc.getNext(k, v, null))
                        if (OperationStatus.SUCCESS != queue.putNoDupData(null, k, v))
                            assert false : "Duplicate email addresses not allowed in queue.";
                }
                else
                {
                    boolean commit = false;
                    Transaction txn = env.beginTransaction(null, null);
                    try
                    {
                        Cursor qc = queue.openCursor(txn, null);
                        try
                        {
                            for (int x = 0; OperationStatus.SUCCESS == uc.getNext(k, v, null);)
                            {
                                if (OperationStatus.SUCCESS != qc.putNoDupData(k, v))
                                    assert false : "Duplicate email addresses not allowed in queue.";
                                commit = true;
                                if (++x == 1000)
                                {
                                    x = 0;
                                    commit = false;
                                    qc.close();
                                    txn.commit();
                                    txn = env.beginTransaction(null, null);
                                    qc = queue.openCursor(txn, null);
                                }
                            }
                        }
                        finally
                        {
                            qc.close();
                        }
                    }
                    finally
                    {
                        if (commit)
                            txn.commit();
                        else
                            txn.abort();
                    }
                }
            }
            finally
            {
                uc.close();
            }
            return this;
        }

        /**
         * Closes the databases and environment.
         */
        public synchronized void close()
        {
            try
            {
                for (Database db : new Database[]{unsent, unsent, recipients})
                    db.close();
            }
            catch (DatabaseException e)
            {
                e.printStackTrace();
            }
            try
            {
                env.close();
            }
            catch (DatabaseException e)
            {
                e.printStackTrace();
            }
        }

        /**
         * Use when populating the databases.
         *
         * @param home The home directory of the blast.
         *
         * @return An environment optimized for single threaded write only access.
         *
         * @throws DatabaseException If there is a problem opening the databases or the database environment.
         */
        public static MyDbI openToFill(File home) throws DatabaseException
        {
            DatabaseConfig dcfg = new DatabaseConfig();
            dcfg.setAllowCreate(true);
            dcfg.setDeferredWrite(true);
            return new MyDbI(home, getFillConfig(), dcfg);
        }

        /**
         * Use for normal blast database access.
         *
         * @param home The home directory of the blast.
         *
         *
         * @throws DatabaseException If there is a problem opening the databases or the database environment.
         */
        public static MyDbI open(File home) throws DatabaseException
        {
            DatabaseConfig config = new DatabaseConfig();
            config.setAllowCreate(true);
            config.setTransactional(true);
            return new MyDbI(home, getDefaultConfig(), config);
        }

        /**
         * @return The number of lock tables based on the number of CPUs.
         */
        public static int getLockTableSize()
        {
            int cpus = Runtime.getRuntime().availableProcessors();
            if (cpus < 4)
                return 1;
            for (cpus = Math.min(MAX_LOCK_TABLES, cpus); cpus > 0 && !PRIMES.contains(cpus);)
                cpus--;
            return Math.max(1, cpus);
        }

        /**
         * Creates a new EnvironmentConfig object suitable for single threaded, write intensive access.
         *
         * @return An EnvironmentConfig object suitable for a single write only thread.
         */
        private static EnvironmentConfig getFillConfig()
        {
            EnvironmentConfig config = getNormalConfig();
            config.setCacheSize(1024 << 10 << 4);
            config.setAllowCreate(true);
            config.setLocking(false);
            return config;
        }

        /**
         * Creates a new com.sleepycat.dbi.EnvironmentConfig object suitable for normal data access patterns.
         *
         * @return An com.sleepycat.dbi.EnvironmentConfig object suitable for normal data access patterns.
         */
        private static EnvironmentConfig getNormalConfig()
        {
            EnvironmentConfig config = new EnvironmentConfig();
            config.setConfigParam("je.log.faultReadSize", "4096");
            config.setConfigParam("je.lock.nLockTables", Integer.toString(getLockTableSize()));
            return config;
        }

        /**
         * Configures an environment for normal transactional access.
         *
         * @return A configuration for normal transactional access.
         */
        private static EnvironmentConfig getDefaultConfig()
        {
            EnvironmentConfig ecfg = getNormalConfig();
            ecfg.setCacheSize(CACHE_SIZE);
            ecfg.setTransactional(true);
            ecfg.setTxnNoSync(true);
            return ecfg;
        }

        static
        {
            PRIMES.add(2);
            PRIMES.add(3);
            PRIMES.add(5);
            PRIMES.add(7);
            PRIMES.add(11);
            PRIMES.add(13);
            PRIMES.add(17);
            PRIMES.add(19);
            PRIMES.add(23);
            PRIMES.add(29);
            PRIMES.add(31);
            PRIMES.add(37);
            PRIMES.add(41);
            PRIMES.add(43);
            PRIMES.add(47);
            PRIMES.add(53);
            PRIMES.add(59);
            PRIMES.add(61);
            PRIMES.add(67);
            PRIMES.add(71);
            PRIMES.add(73);
            PRIMES.add(79);
            PRIMES.add(83);
            PRIMES.add(89);
            PRIMES.add(97);
            PRIMES.add(101);
            PRIMES.add(103);
            PRIMES.add(107);
            PRIMES.add(109);
            PRIMES.add(113);
            PRIMES.add(127);
            PRIMES.add(131);
            PRIMES.add(137);
            PRIMES.add(139);
            PRIMES.add(149);
            PRIMES.add(151);
            PRIMES.add(157);
            PRIMES.add(163);
            PRIMES.add(167);
            PRIMES.add(173);
            PRIMES.add(179);
            PRIMES.add(181);
            PRIMES.add(191);
            PRIMES.add(193);
            PRIMES.add(197);
            PRIMES.add(199);
            PRIMES.add(229);
            PRIMES.add(241);
            PRIMES.add(241);
            PRIMES.add(241);
            PRIMES.add(271);
            PRIMES.add(283);
            PRIMES.add(283);
            PRIMES.add(313);
            PRIMES.add(313);
            PRIMES.add(313);
            PRIMES.add(349);
            PRIMES.add(349);
            PRIMES.add(349);
            PRIMES.add(349);
            PRIMES.add(421);
            PRIMES.add(433);
            PRIMES.add(463);
            PRIMES.add(463);
            PRIMES.add(463);
            PRIMES.add(523);
            PRIMES.trim();
        }
    }
}

Saturday, September 15, 2007

The Missing Minute

Sometime earlier today
A minute of mine went away
So I thought about main
But that thought was in vain
The minute was hiding okay?

I opened the profiler quick
And attach to the app in a nick
I collected a sample
It wasn't quite ample
The minute continued to tick.

I decided to give it a rest
In order to give it my best
I went for a drive
And in about five
My mind was finally unstressed.

But on the way home I found it
The place where the minute was grounded
It was caught in a latch
With a timer dispatch
But no signal handler around it.

Let this be a lesson to all
When debugging processes stall
Take a step back and maybe a nap
And the bug is likely to fall.

Thursday, July 12, 2007

Becoming a concurrency expert. Rule number not one, Relax.

The holy grail for a concurrency expert is wait-free, and if you can't achieve that, lock-free code because generally they are more scalable than algorithms that use locks. The previous link and this one describes some of the benefits of non-blocking synchronization but I'll give you a contrived example.

LF/WF algorithms can't deadlock! No locks, no deadlocks. Deadlocks are a bane to scalability because in the wild they are probabilistic events whose probability increases as concurrency increases. So you can imagine a situation where you have some code that runs smoothly on your old single processor system, then you upgrade to a dual core system and suddenly the code starts freezing every once in a while. You think to yourself, "gremlins" and continue along your merry way. Christmas comes early and you win a shiny new quad core box from some tech event in Atlanta. So you are thinking, "4x the processing power, 4x the performance, yeah!" But instead your program is freezing all the time. Simply killing it and restarting is not good enough anymore. You have a deadlock on your hands. You've gone from a single processor to a 4-way system and the scalability of the program has not followed suit. This is not atypical of a lot of Java programs in the wild, because though it may come as a shock to you (it sure as hell shocked me!), most Java programmers don't know spit about concurrency, even though Java has concurrency baked in from birth! And now that multi-core systems are becoming the norm, concurrency bugs that have laid dormant for years are waking up and stinging users.

It's possible, with a great deal of effort, to eliminate deadlocks in non LF/WF code. So (a) given that writing LF/WF code is as hard or harder than writing deadlock free code and (b) you are willing to solve any deadlock problems in the non LF/WF code, is LF/WF worth it? As with all questions relating to trade-offs, the answer is "it depends". In this case, it depends on how much throughput is enough. In the WF case you are guaranteed system-wide throughput with starvation freedom while in the LF case you are still guaranteed system-wide throughput but with the possibility that individual threads may starve. The bottom line is, progress is always being made. There is no such guarantee with blocking synchronization.

Because of the complexity associated with LF/WF algorithms most programmers never tackle LF/WF head on. Contact with LF/WF algorithms and code come in the form of using LF/WF datastructures (i.e. java.util.concurrent.ConcurrentLinkedQueue). But it may surprise you that in your own code there may be opportunities to write LF/WF code.

Disclaimer:

I'm not advocating everybody going through every line of code and trying to make it LF/WF (though I do advocate going through every line of your code and making sure it's thread safe). You really, really, need to have an extremely strong grasp of the Java Memory Model before you can even begin to think about writing LF/WF code, especially the happens-before rules.

You should limit your LF/WF tinkering to critical paths only. Critical paths are hot sections of code (code that is executed frequently). You need two tools to fix critical paths. Firstly, you need a Java profiler to tell you where the critical path is. The critical path is going to be the method call chain where the program spends the majority of it's time while under load. The under load distinction is extremely important because if you take a server application as an example, and profile it when there is no load, the profiler is going to report that the app is spending most of its time in [something like] Socket.accept(), which doesn't tell you anything about the performance of the app. In your quest for the critical path the best any Java profiler can do is tell you what methods are consuming the most amount of time. They cannot peer into the method and tell you which specific line of code is slow or if a lock is hot (a hot lock is one that is highly contended). This is where the second tool comes into play. You need a hardware profiler.

A hardware profiler differs from a Java profiler in that it show events at the CPU. Every modern CPU comes with all sorts of counters that enables programs to know what's going on inside the CPU. It can tell you things like cache hit/miss rates, stalls, lock acquisitions and releases, etc. Some operating systems comes with hardware profilers baked right in. Solaris 10/OpenSolaris on Sparc is the gold standard when it comes to observability. mpstat, corestat, plockstat, and [the big daddy of them all] DTrace are some of the tools baked into Solaris 10/OpenSolaris that allow you to dig deep into the bowels of the system to figure out exactly what's going on. If you aren't running Solaris but Linux or Windows on AMD you can use AMD's CodeAnalyst Performance Analyzer. Finally, if you are running Linux or Solaris (Intel or AMD) you can use Sun Studio 12 to get at the data. All the hardware profiler tools I've mentioned are free and/or open source. So you have no excuse not to have at least one installed.

So here are the steps you've completed so far:

  1. Profiled the app under load.
  2. Identified the critical methods (hotspots).
  3. Tuned the critical method(s) as best you can.
  4. Repeat steps 1-3 until you hit the point of diminishing returns.
At this point if the throughput is where you want it you can stop. You didn't have to write a lick of LF/WF code. Good for you. But what happens if you want to see how far you can really push the system? You start really cranking up the amount of threads (assuming you have available processors to execute them concurrently). You take at the look at the system's CPU utilization and it's redlining. You need to make sure it's actually doing useful work. So now it's time to fire up the hardware profiler to see what's going on in the CPU.

Aside:

Sun Studio 12 does the best job, of the tools listed, of associating CPU events back to the source code line(s) that produced them.

So you fire it up and the first thing that jumps out at you is you have a smokin' hot lock along your critical path. If you can't reduce the granularity of the lock or its scope, going LF/WF may be your only option.

Aside:

It's entirely possible that you can neither reduce the granularity of the lock nor rewrite the critical section as LF/WF. At this point you are screwed. You are just going to have to buy another box (or virtualize) and spread the load. If you can't spread the load you are royally screwed so the best thing to do is degrade gracefully.

I've gone through all of that setup just so I dump some code on you:

public class RecipientLimits
{
    private static final long ONE_HOUR_MILLIS = TimeUnit.HOURS.toMillis(1);

    private static final long THREE_DAYS_MILLIS = TimeUnit.DAYS.toMillis(3);

    private static final long ONE_MONTH_MILLIS = TimeUnit.DAYS.toMillis(31);

    private final ConcurrentMap hosts = new ConcurrentHashMap();

    /**
     * Tracks the last time we pruned the table.
     */
    private final AtomicLong lastrun = new AtomicLong(System.currentTimeMillis());

    /**
     * Get the recipient limit for host.
     *
     * @param host The host.
     *
     * @return The maximum number of recipients the host will accept.
     */
    public int get(InetAddress host)
    {
        Entry e = hosts.get(host);
        if (null == e)
        {
            Entry t = hosts.putIfAbsent(host, e = get(host.getAddress()));
            if (null != t)
                e = t;
        }
        return e.limit();
    }

    /**
     * Confirms that host accepts limit recipients.
     *
     * @param host  The host.
     * @param limit The number of recipients that was accepted.
     */
    public void confirmed(InetAddress host, int limit)
    {
        Entry e = hosts.get(host);
        if (null != e)
            e.confirm(limit);
        prune();
    }

    /**
     * Indicates that host did not accept limit number of recipients.
     *
     * @param host  The host.
     * @param limit The limit.
     */
    public void denied(InetAddress host, int limit)
    {
        Entry e = hosts.get(host);
        if (null != e)
            e.decrement(limit);
        prune();
    }

    /**
     * Removes inactive hosts from the table.
     */
    private void prune()
    {
        long last = lastrun.get();
        long millis = System.currentTimeMillis();
        if (millis - last >= ONE_HOUR_MILLIS && lastrun.compareAndSet(last, millis))
        {
            for (Iterator i = hosts.values().iterator(); i.hasNext();)
            {
                Entry e = i.next();
                if (0 == e.users.get() && millis - e.lastaccessed.get() >= THREE_DAYS_MILLIS)
                    i.remove();
            }
        }
    }

    /**
     * Look up host in the database.
     *
     * @param host The ip address of the host.
     *
     * @return A new Entry.
     */
    private Entry get(byte[] host)
    {
        //@todo don't forget to create an entry in the database if host does not already exist.
        return new Entry(1);
    }

    final class Entry
    {
        /**
         * The last time (milliseconds timestamp) this enty was accessed.
         */
        final AtomicLong lastaccessed;

        /**
         * The number of threads reading/writing this object.
         */
        final AtomicInteger users = new AtomicInteger();

        /**
         * Semaphore for updates.
         */
        private final AtomicInteger dflag = new AtomicInteger();

        /**
         * The recipient limit.
         */
        private final AtomicInteger limit;

        /**
         * Indicates when we've maxed out {@link #limit}.
         */
        private volatile boolean maxo;

        /**
         * The millisecond timestamp when we maxed out date.
         * 

* This field is correctly synchronized because there is a happens-before edge created by the write to this * followed by a write to {@link #maxo} in {@link #decrement(int)} and then the read of {@link #maxo} followed by the read * of this in {@link #confirm(int)}. *

*/ private long maxtstamp; /** * Constructs a new Entry. * * @param limit The recipient limit. */ Entry(int limit) { this.limit = new AtomicInteger(limit); lastaccessed = new AtomicLong(System.currentTimeMillis()); } /** * @return The current limit. */ int limit() { users.incrementAndGet(); try { return limit.get(); } finally { lastaccessed.compareAndSet(lastaccessed.get(), System.currentTimeMillis()); users.decrementAndGet(); } } /** * Confirm the limit. * * @param limit The value returned by {@link #limit()}. */ void confirm(int limit) { users.incrementAndGet(); try { if (0 == dflag.get() && (!maxo || System.currentTimeMillis() - maxtstamp >= ONE_MONTH_MILLIS) && this.limit.compareAndSet(limit, limit + 1)) { //@todo - talk to database //@todo - use limit not limit + 1 } } finally { lastaccessed.compareAndSet(lastaccessed.get(), System.currentTimeMillis()); users.decrementAndGet(); } } /** * Decrement the limit. * * @param x The value returned by {@link #limit()}. */ void decrement(int x) { users.incrementAndGet(); try { if (x > 1 && dflag.compareAndSet(0, x)) { try { if (limit.compareAndSet(x, x = x - 1)) { boolean dbput = true; if (dbput) { maxtstamp = System.currentTimeMillis(); maxo = true; } } } finally { dflag.set(0); } } } finally { lastaccessed.compareAndSet(lastaccessed.get(), System.currentTimeMillis()); users.decrementAndGet(); } } } }

Let's revisit the title of this post "Becoming a concurrency expert. Rule number not one, Relax". I've emphasized relax because it is critically important to finding LF/WF opportunities in your own code. So what exactly does relaxing mean? The biggest thing it entails is realizing that you have very little control over the order in which threads execute and being OK with that. Because if you try to be draconian about the order in which things happen you will have to go single threaded or use locks. So once you've relaxed and let go of draconian ordering the only thing you have to worry about is ensuring that data races are benign. Notice I didn't say eliminate data races, I said "ensuring that data races are benign". The difference of course is relaxation. In LF/WF, data races are part of the design because at some point in the code you are going to need to do a CAS (i.e. java.util.concurrent.atomic.AtomicBoolean.compareAndSet(false,true), java.util.concurrent.atomic.AtomicInteger.incrementeAndGet(), etc) and a CAS is a [CPU supported] race condition waiting to happen. CAS isn't the only place where its okay to let a race condition go unchallenged. Anywhere you can prove that the consequence of a data race is benign is an opportunity to relax.

Let me back up quickly and talk briefly about ordering. I hope you don't think I said that ordering is not important or that you have absolutely no control because that is not true. Let me repeat it. What you don't have control of is when a thread will run. That's [ultimately] the responsibility of the operating system. What that translates into is, you don't have control of when something executes, only what executes.

The new JMM strengthened the guarantees of volatile to prevent the reordering of volatile read/writes in relationship to non volatile fields. In other words, you can use volatile fields to force code to execute in a certain order without the use of the synchronized keyword. Which also means you can use a single volatile field to safely publish multiple non volatile fields (an example of this is located in the code above and is described in the JavaDoc comment for maxtstamp). This was not possible prior to Java 5. Before Java 5 if you wanted visibility guarantees for fields you had to (a) declare all of them volatile or (b) use a synchronized block.

So there is a lot you can accomplish given the new volatile semantics but you can't do everything. Specifically, you can't make binding decisions with volatiles. So what's a binding decision?

if (some_volatile_condition IS True)
{
  //Execute code under the assumption that some_volatile_condition is still True.
}
... is a binding decision and in the absence of locking [before the read of some_volatile_condition] is a bug, except of course, in the case that the code being executed results in a benign race condition. Examples of non-binding decisions can be found in Entry.confirm(int) and Entry.decrement(int).

That's it for today. I'll pick the code apart in Part II. Thanx for stopping bye.

Tuesday, June 26, 2007

A Developer's Journal: Solaris #11

CTRL+BACKSPACE has bitten me for the last time.

I've been trying to make OpenSolaris/JDS my work environment for the last week and time and time again my desktop restarts itself because of CTRL+BACKSPACE. 99% of the time I swear I didn't press the key combination! Fixing it has solved one of the many things about JDS that annoy the crap out of me (I'll tell those stories another day).

The solution is extremely simple. Edit the xorg config file located at /etc/X11/xorg.conf and add the following line to the ServerFlags Section:

Option         "DontZap"  "true"
Save the file and hit CTRL+BACKSPACE to restart the JDS for the last time.

I'm so elated I figure I would take a moment and send this one up. Thanx Alfred Peng.

Monday, June 18, 2007

A Developer's Journal: Solaris #10

I've solved the font problem. It was a font rendering issue. I have dual 20 inch LCD monitors and the default font rendering settings just wasn't cutting it. A quick visit to the "Font" dialog box under "Preferences" made all the difference. The fuzz is gone. Yeah!

I've also completed configuring my bash environment ala Gentoo so Solaris is starting to feel a lot less foreign to me than it did two days ago. The next step is installing and configuring IDEA and pulling down my current working set from a couple Subversion repositories. I don't think the real benefits of Solaris are going to be evident to me without hacking on some code.

Saturday, June 16, 2007

A Developer's Journal: Solaris #9

Installation complete! Yeah!

Initial Reactions

The fonts look like crap! At this point, I'm not sure if the problem is the default set of fonts or if font rendering just sucks. Every bit of text everywhere is fuzzy. Text is my life! I am a programmer and sometimes blogger, after all. Solving this issue is at the top of my todo list.

The default configuration for the root user account is just plain insipid. The root user does not have a default HOME directory so every file and directory that gets automatically created by the shell or the desktop environment gets dumped to /. How absolutely, positively, retarded is that? Every Linux distro I have ever used creates a /root directory to store the root user's files. I just can't fathom the rationale behind not setting a decent default directory for root.

Now some of you may say or be thinking, "you are not supposed to be using the root user account anyway". If you did say that then you've obviously never installed Solaris 10. Because if you have installed it then you know you only ever get prompted to provide a password for the root user. The install process forces you to wait until the installation is complete to create a normal user account. The problem is on first reboot the only account you have to log in with is root and the minute you log in, the desktop environment creates a host of hidden files and directories on the root filesystem for root. So you never get the opportunity not to use the root account. Like I said before, insipid, and that's putting it nicely.

The good news is most of my hardware works out of the box. The only thing that isn't working is my USB DVD-R/W. One more item for my todo list.

A Developer's Journal: Solaris #8

I am returning to the world of Solaris. My original foray into the land of (the) Sun didn't go so well. It was a try and buy 60 day trial of a T1000 that just didn't go a well as I had hoped. Some of the problems were Sun's fault but most of them where mine. One of the big issues was proximity. The T1000 was hosted in an off site data center which required several layers of indirection to get to the machine and some things are just harder to do remotely. The bottom line is I didn't get the most out of the 60 days. But I really am interested in learning more about Solaris. DTrace is simply too compelling a technology to ignore, especially now that Solaris is opensource. So almost 8 months later I'm revisiting the land of Sun.

This time around I'm doing things a lot closer to home. I've added a SATA RAID controller and 4 35GB 10,000 RPM HD to my workstation. This will be Solaris' new home. I had to get rid of my CD-R/W for a 3 disk enclosure but I have an external USB DVD-R/W so no sweat.

A standard Solaris install is bit more heavy weight that I'm interested in. Coming from the land of Gentoo, I'm used to and prefer installing things as I need them instead of a default kitchen sink install like the one that comes w/ standard Solaris 10. Plus, I'm a tinkerer. Fortunately for me, Sun saw me and my kind coming and has a distro taylor made for us; OpenSolaris Developer Edition. We get the latest and greatest Solaris kernel and the developer tools necessary to tinker, compile, and profile. It has a (relatively) recent GNOME based desktop and Firefox is the default browser.

As soon as I'm done posting this I'm going to go burn my (downloading) distro to DVD and install it. Stay tuned.

Thursday, May 10, 2007

I'm A Real Programmer Because ...

David Miller says I'm a real programmer. Are you?

Wednesday, March 28, 2007

A Great Hot Dog

Hebrew National Franks really does make the best hot dogs I've ever eaten. At this point if it ain't Hebrew National I ain't eatin' it. After all, it's kosher. So if it's good enough for Jesus it's good enough for me. But then again Jesus was something of a radical. Maybe he bucked Jewish law and ate non kosher beef. If he did and was trying to spite the establishment and was still around, he would be missing out on a great hot dog.

If you are going to have a great hot dog you are going to need a great hot dog bun. Otherwise, what's the point? The best hot bug bun I've found is Martin's Long Roll Potato Rolls and according to the packaging, it has that "famous Dutch taste".

My condiment choices are simple yet cosmopolitan. It requires mayonnaise, pesto, sharp or extra sharp cheddar cheeses and ketchup (optional). Hellmann's is the only brand of mayonnaise I'll buy. As they say, "if you bring out the Hellman's you bring out the best" [Those TV people ain't never lied!]. I've tried different brands of pesto and can't say I've found a favorite brand. They all bring something a little different. I'm currently using Classico brand pesto, in case you want a place to start. Cracker Barrel shredded sharp cheddar is my cheddar of choice. I also like Sargento but Cracker Barrel has just a little more bite to it. Finally, we get to ketchup. I'm a Heinz man. Been that way since I was a boy. I'll try others when outside the domicile and Heinz is not available, but never while I'm at home.

So now its time to put it all together. Boil the hot dogs, cover the pot, turn the stove off, and immediately ...

  1. Spread mayo on half of the bun.
  2. Spread pesto on the other half.
  3. Lay down a light layer of cheese in the bun groove.
  4. Drop a hot dog in the groove.
  5. Add another layer of cheese on top.
  6. Let sit for 30-60 seconds.
  7. Enjoy!
Let me expand step 6 briefly. The reason there is a step 6 is to allow the heat from the cooling hot dog to lightly melt the (refrigerated) cheese.

You may have noticed that the application of ketchup was not in any of the steps. That's because you have a choice to make. It's been my observation that the ketchup weakens the impact of the cheese on the pallet. Thus, if you are afflicted by this you may want to skip the ketchup entirely. But since one hot dog is never enough, make two, then you can have it both ways.



P.S. I'm not Jewish. Not even a little bit.

Saturday, November 18, 2006

Getting IDEAs Flowing

This is a story about starting IDEA 6 on my Gentoo GNU/Linux machine. IDEA 6 has been out a while now, but I'm not the kind of person to buy the first release of a new version of software. I like to wait until the more adventurous types take it for a spin and shake the first few waves of bugs out. So now that it's at maintenace release 2 I'm taking it out for a spin.

Getting it running was a bitch. With the default settings I could get as far as the "loading project" dialog box. It just hung there, forever, which was long enough for me to notice that they've fixed one of (a few) annoyances I have with version 5.

Aside:

Whenever IDEA is running a command that is taking a while, the "tip of the day" dialog box attaches itself to the progress meter. In 5 the dialog box is vertically too small to read the text inside it. I always have to pull the edge of the box down to see the tip. After the first few times I gave up on the "tip of the day" all together. After I realized I can actually see what's in the box I started flipping through them while I waited for IDEA to complete startup. Very insightful stuff. It turns out there are a plethora of shortcuts that I didn't know about. Which sucks. Think off all the time I could have saved w/ the new shortcuts instead of mousing around the editor.

At first I thought it was just busy importing and converting the settings and configuration from the old install. So I fired up Spider Solitaire in XP/VMWare and started listening to episode 236 of This American Life. That's how I normally pass the time when I need to wait for a computer to do something and I don't want to change my mindset to a different computer related task.

After fifteen minutes it still wasn't done and trying to close the window was getting me nowhere. So I murdered it with a 9 caliber signal. Fortunately I had started it from a shell so I could see the STDERR messages. The JVM was throwing OOM errors related to a too small perm gen size. A quick cat of {INSTALL_DIR}/bin/idea.vmoptions revealed a mandatory limit of 99MB. An easy fix. Just remove the offending line because the JVM --Sun JVM 1.5_08 Linux, to be specific-- defaults to ~130MB. Fired it up again and got to the main menu. Yippee!! Things are looking up.

Wrong!

I tried opening the last project I was working on and the project loads but the editor refuses to accept input. Then the window loses focus and one of the CPUs hit 100%. top tells me IDEA is the offending process. I figured I would give it more time. Maybe it will settle down after a while plus I've got This American Life on pause and episode 236 is quite good. So I unpause and started another round of Spider Solitaire. A few minutes go by and the JVM starts complaining the perm gen size is still too small. No problem. I run a server class machine w/ 12GB of RAM of which 4GB is in use so I've got plenty to spare. So I run echo "-XX:MaxPermSize=256m" >> {INSTALL_DIR}/bin/idea.vmoptions and rerun the start script. It gets to the main menu in a reasonable amount of time. So far so good. I reopen the project and cross my fingers. Booyah! 256MB is the magic number. Hooray!!

Friday, October 06, 2006

A Developer's Journal: Solaris/CMT #7

I got the answers I was looking for:

Solaris has major releases every 2 - 3 years and minor releases every 3 - 6 months S10 is the current major release.

...

The updates have major features integrated. To move from Update to Update you usually re-install the OS. You download the CD image or put it on your install server. You can opt for a Upgrade option as well. In between updates there are patches. These fix problems with a variety of products. There is one patch in particular called the kernel jumbo that has many of the fixes for just the OS. These can be downloaded from sun.com. There are mandatory patches to fix panics etc. These are available for free download. Others require a support contract.

A Developer's Journal: Solaris/CMT #6

I'm back in Solaris mode and I've hit a wall. Solaris 10 is Unix and that's great because visa-vi GNU/Linux I am very comfortable with Unix. Before I got my hands on my box I made the decision not for force my GNU/Linux habits onto Solaris because its probably a recipe for pointless frustration. So I'm very interested in learning the Solaris way of doing things. The issue I'm struggling with right now is patching and versioning. I don't understand how they fit together as a whole. Now I know I just said I wasn't going to force my GNU/Linux habits onto Solaris but I need to do a little compare and contrast to ensure that I'm asking the right question.

In Gentoo there is no explicit concept of a version. There are live CD releases such as 2006.1 or 2006.2 but you can't use that as a mechanism for discussing the version of Gentoo that's actually running on your box because as soon as you run emerge --sync && emerge -uD world or build your own kernel the whole version concept breaks down.

At the other extreme is Windows. You buy a copy of Windows of a specific version and it stays that way until Microsoft releases a new version and you explicitly purchase and install the new version. In the interim the most you can do is apply hot fixes and service packs. So to move from WinXP to Vista I would have to buy Vista and install clean or do an upgrade of my current WinXP install.

So which one is Solaris most like? Gentoo or Windows? Because I'm trying to get from Solaris 10 1/06 to 6/06 and I'm not sure how to get there.

Thursday, October 05, 2006

Throw more threads at it

Why Events Are A Bad Idea (for high-concurrency servers)

An interesting position paper on why parallel programming using threads are better for highly concurrent servers compared to event based approaches. Even though the paper is 3+ years old its more applicable now than it was then in large part because threads have gotten cheaper. How so?

  • High performance thread libraries are part of the Linux kernel since 2.6 (NPTL) and the Solaris 10 kernel both of whom implement a 1x1 model.
  • As memory performance and bandwidth increases the cost of context switches decreases. As a matter of fact reducing the cost of a context switch is baked into the DNA of Sun's Niagara family of processors. AMD had the bright (right) idea to include the memory controller right on the chip. They also switched to a NUMA topology which boosts performance when the data is closer to where it's needed. Intel and IBM continue to increase on chip cache sizes to reduce the frequency and thus the latency associated with having to fetch data from main memory.
  • Massively parallel super computer on a chip is around the corner. Check out:

I could go on but I'll stop here. Click on the link, read the paper, and draw your own conclusions.

Sunday, October 01, 2006

OpenBSD needs our help

Dear Reader,

I hope to appeal to your sense of decency and fairness in requesting that you support OpenBSD's ongoing campaign to get Intel to provide open access to the documentation (and in some cases binary blobs) behind their wireless chipsets [full story here]. This access will not only benefit OpenBSD but all FLOSS operating systems and strengthen the ecosystem by giving us (you and me) more choices from a hardware and operating system perspective.

Please stop buying Intel's products if possible and at a minimum send emails to majid.awad@intel.com and voice your concerns. The Internet gives you a voice and I implore you to use it.

Sincerely,


HashiDiKo

Saturday, September 30, 2006

Java Keeps Keeping On

The real motives why industry analysts love to predict Java fall... and why they will spend another 10+ years being proven wrong.

A very good entry discussing some of the reasons why Java continues to defy the doomsday predictions of so-called "industry experts".

Java not being the new kid on the block is not news it's an achievement that every professional (state of mind not state of employment) Java programmer can be proud of.

Wednesday, September 13, 2006

Concurrent Javadoc

Writing correct concurrent code is hard and what I'm finding equally challenging is documenting concurrent behavior in code in prose (javadoc to be specific).

Pre Java 1.5 all we had where java.lang.Thread, the Runnable interface, the synchronized keyword, and volatile variables. Because synchronized was the only real mechanism for concurrency control if you wanted to reason about the thread safety of a class all you had to do was look for the synchronized keyword. Some of you may be asking "but what about volatile?" volatile then was nowhere near as powerful as it is now. Back then the only thing volatile really meant was do not cache. Nowadays volatile not only means do not cache it also affects the visibility of other non volatile variables and it also imposes ordering constraints. Basically synchronized and volatile no longer act in isolation. All activities inside the JVM is now governed by a well defined memory model (JMM). It's no longer enough to look for the synchronized keyword you have to keep the memory model in mind at all times which is damn hard to document. Therefore I'm seeking feedback.

Imagine you have been asked to maintain a body of code with which you have had no previous experience and the only thing you know is it's targeted at Java 1.5 and higher (new JMM). You come across this (partial) class definition:

1    /**
2     * Ties {@link BlastMaster}, {@link BlastManager}, and {@link BlastAgent} objects together around a blast. It is the
3     * highest layer of abstraction available on a blast. It provides the hooks for the lifecycle management of a blast and
4     * is the communication channel between BlastAgent(s) and BlastManager(s) and BlastMaster.
5     * <p>
6     * The lifecycle of a context goes something like this. It is created by BlastMaster, injected into the outbound queue
7     * by BlastMaster, scheduled by BlastManager, delivered by BlastAgent(s), descheduled by BlastManager, rescheduled by
8     * BlastManager, delivered by BlastAgent(s) ....
9     *
10    * The scheduling, delivering, and descheduling steps repeat until a) all the recipients have been sent the blast or b)
11    * the duration of the blast has expired.
12    * </p>
13    *
14    * @author HashiDiKo
15    */
16   public final class BlastContext extends Constants
17   {
18       /**
19        * Keeps track of the number of domains yet to be delivered. When this number reaches zero it means the current
20        * iteration of trying to deliver the blast has ended.
21        *
22        * @see BlastAgent#decrementRemaining()
23        * @see #fire(ContextEvent.Event)
24        */
25       final AtomicInteger remaining = new AtomicInteger();
26   
27       /**
28        * Provides a mechanism for a {@link BlastManager} to be notified when the {@link BlastAgent}s executing the context
29        * stop executing it.
30        * <p> This object does a great deal in terms of establishing <em>happens-before</em> edges between a blast manager
31        * and its agents and between blast agents.</p>
32        */
33       final Set<Thread> agents = Collections.synchronizedSet( new HashSet<Thread>( MAX_BLAST_AGENTS + 1) );
34   
35       /**
36        * The message.
37        *
38        * <h4>Visibility</h4>
39        *
40        * <p>The blast manager thread is the only thread that writes this field while blast agents are the only threads
41        * to read this field. This field is <i>correctly synchronized</i> because the blast master writing this field
42        * <b>always</b> <i>happens-before</i> the blast agents reading it. This <i>happens-before</i> relationship is
43        * established in the code starting with {@link #init()} and its call to {@link #doInit()}. {@link #doInit()} is
44        * the only place in the code where this field is written. You simply need to follow the flows in
45        * {@link BlastManager} starting from where {@link #init()} is called to see synchronization the points that
46        * establishes the <i>happens-before</i> relationship between the blast manager and its blast agents in regard to
47        * this field.
48        * </p>
49        */
50       Message message;
51   
52       /**
53        * Indicates that this context is context to an old blast. In other words the system was restarted and the blast
54        * was loaded from the database.
55        *
56        * @see #init()
57        * @see #BlastContext(boolean, EmailBlast, BlastMaster)
58        */
59       private final boolean old;
60   
61       /**
62        * The home directory of the blast.
63        */
64       private final File home;
65   
66       /**
67        * The duration of the blast in milliseconds.
68        */
69       private final long durationMillis;
70   
71       /**
72        * Used to implement our backoff algorithm.
73        *
74        * <h4>Visibility</h4>
75        *
76        * <p>The field is only ever read/written by blast agents thus the visibility of this field is limited to
77        * blast agents in general and specifically the {@link #fire(ContextEvent.Event) fire} method. Even though
78        * {@link #fire(ContextEvent.Event) fire } is not <tt>synchronized</tt> and this field is non volatile it is
79        * <em>correctly synchronized</em>. There are 2 things at work that make this so. Firstly, there is never
80        * concurrent access to this field. At most 1 blast agent thread will write/read this field for a delivery cycle.
81        * This is guranteed by {@link BlastAgent#decrementRemaining()} because a) how its implemented and b)it is the only
82        * place where a blast agent get accesses this field. Finally, the fact that blast agents call
83        * {@link #register(Thread)} before they execute a blast and {@link #unregister(Thread)} when they are done
84        * executing a blast (both contained <tt>synchronized</tt> blocks using the same monitor) gurantees that the update
85        * of this field will be visible to any subsequent reads by different blast agents.
86        * </p>
87        * @see #fire(ContextEvent.Event)
88        */
89       private int magnitude;
90   
91       /**
92        * Used to implement our backoff algorithm.
93        *
94        * <h4>Visibility</h4>
95        *
96        * <p>The field is only ever read/written by blast agents thus the visibility of this field is limited to
97        * blast agents in general and specifically the {@link #fire(ContextEvent.Event) fire} method. Even though
98        * {@link #fire(ContextEvent.Event) fire } is not <tt>synchronized</tt> and this field is non volatile it is
99        * <em>correctly synchronized</em>. There are 2 things at work that make this so. Firstly, there is never
100       * concurrent access to this field. At most 1 blast agent thread will write/read this field for a delivery cycle.
101       * This is guranteed by {@link BlastAgent#decrementRemaining()} because a) how its implemented and b)it is the only
102       * place where a blast agent get accesses this field. Finally, the fact that blast agents call
103       * {@link #register(Thread)} before they execute a blast and {@link #unregister(Thread)} when they are done
104       * executing a blast (both contained <tt>synchronized</tt> blocks using the same monitor) gurantees that the update
105       * of this field will be visible to any subsequent reads by different blast agents.
106       * </p>
107       * @see #fire(ContextEvent.Event)
108       */
109      private boolean hrs;
110  
111      /**
112       * Stores the backoff timeout (the end result of the backoff algorithm computation).
113       *
114       * <h4>Visibility</h4>
115       *
116       * <p>The visibility requirements for this field is kind of funky. It's is a read/write field for
117       * {@link BlastManager}s and write only field for {@link BlastAgent}s. So the core visibilty issue is writes by
118       * blast agents being visible to blast manager. The <em>happens-before</em> edge that makes this possible is
119       * established by the monitor acquisition of {@link #agents} by the blast agent thread that writes this field and
120       * the subsequent acquisition of the same monitor by the blast manager.
121       * </p>
122       *
123       * @see #fire(ContextEvent.Event)
124       */
125      private long backoffTimeout;
126  
127      /**
128       * Indicates if {@link #doInit()} has been called.
129       * This field is only ever read/written by the blast manager thread for this context thus it is correctly
130       * synchronized.
131       */
132      private boolean didInit;
133  
134      /**
135       * Temporarily store the short domain blacklist expiration.
136       * <p>
137       * When this value is greater than zero it indicates to the blast manager thread that even though {@link #domainQ}
138       * is empty this context has undelivered recipients.
139       * </p>
140       * <p>This field is <em>correctly synchronized</em> because it is only ever read/written by the blast manager
141       * thread.</p>
142       *
143       * @see #getDomains()
144       */
145      private long shortestDomainTimeout;
146  
147      /**
148       * Gate to {@link #manager}. It establishes a <em>happens-before</em> edge with {@link #manager}.
149       */
150      private final CountDownLatch managerLatch = new CountDownLatch( 1 );
151  }
152  

So. Does the javadoc comments say enough about how non volatile class members manage to be correctly synchronized?

Saturday, September 09, 2006

A Developer's Journal; Solaris/CMT #5

I am feeling disingenuous about my "A Developer's Journal; Solaris 10 ..." title because I've made 5 entries so far and have said very little about Solaris 10. The truth of the matter is Solaris is the means and not the end. Let me explain. I have absolutely no interest in Solaris as an end to itself at this point because of Linux. And though Sun has made awesome progress with Solaris on x86 it still has a way to go before it is as seamless as Solaris on Sparc. So my key interest in Solaris is strictly with Solaris on Sparc in general and Solaris on Niagara processors specifically. Therefore all entries, starting with this one, will be titled "A Developer's Journal; Solaris/CMT ...". My apologies to any readers who have arrived here under the assumption that these are strictly Solaris 10 entries.

Friday, September 08, 2006

A Developer's Journal; Solaris 10 #4

I brought it home. It's still in the box. It's setting in my living room steering at me and refuses to blink. I know I don't stand a chance of winning the steering game but I steer anyway eyes bloodshot I steer. It's all I can do with it at the moment. I have deadlines to meet and missing deadlines is bad for business so I can't cut into my dev time to play with the it. The clock is ticking, ticking, ticking. As I look away from my monitor to rest my eyes it catches my gaze. It steers at me and I steer back neither of us blinking while the clock ticks, ticks, ticks.

Monday, September 04, 2006

A Developer's Journal; Solaris 10 #3

The T1 is here. Andrew called to let me know it's at the office. It only took 6 days. I'm impressed. I guess Sun has finally got the kinks in there supply chain hammered out.

I need to figure out where I'm going to put it. I can set it up at the office or take it down to the data center or set it up at home. I just don't know right now because my thoughts are elsewhere. I've got some integration work to do with Jetty and a heck of a lot of code to write for the web service interface for my project.

So much to do so little time.

Sunday, September 03, 2006

Latches in Java

David Holmes provides an excellent description of the difference between java.util.concurrent.CountDownLatch and java.util.concurrent.CyclicBarrier on the concurrency interest mailing list. I have posted it here for those of you who don't subscribe to the list and may not have a good grasp of these classes. And since I use CountDownLatch in code I've posted in the past and in code I'm planning on posting I'm happy to have such a clear explanation that I can link to in future entries.

... As Brian describes one notion of state "latching" is that it progresses to a terminal value from whence it no longer changes. It is permanently "latched". A synchronization object that behaves in this way is the CountDownLatch - once "open" it never closes and can't be reset.

More generally though latches can be reset - consider a digital flip-flop such as a "gated D-latch", the flip-flop latches the value of the data when the gate is pulsed/strobed; if you change the data without changing the gate then the latched value is unchanged, but change the gate and the latched value is updated.

In synchronization object terms a latch is sometimes called a "gate" - the connotation being that if the gate is open anyone/everyone can pass through; while if it is shut no one can pass through. The CountDownLatch operates this way, but a more general "gate" is the CyclicBarrier which can also be reset (and automatically does so). Of course the semantics of CountDownLatch and CyclicBarrier are somewhat different. CyclicBarrier is what can be called a "weighted gate" or "weighted bucket" - it is set up to expect N threads to arrive, when they arrive they have sufficient "weight" to open the gate, or tip the bucket - in this case the gate/bucket is spring-loaded and closes/rights-itself as soon as the threads leave, so it is ready for the next set of threads to use. CountDownLatch on the other hand is like a gate with multiple padlocks - when the last padlock is removed, the gate opens and it stays open. Aren't these analogies quaint :-) We could have defined CountDownLatch to allow reset but reset semantics are messy and usually not needed for CDL usage, in contrast barrier designs typically always use the barrier multiple times.

It seems the database folk are using the term "latch" for lightweight lock, which is an uncommon usage from a synchronization perspective and a poor choice in my view, though arguably there is an analogy between "locking a door/window" and just "latching it shut". In that sense "latching" is a weaker form of "locking". But I don't like the usage.

A Developer's Journal; Solaris 10 #2

After 2 "Try and Buy" applications Sun Microsystems has finally shipped my T1. The thought of getting my hands on it gets my geek johnson tingling. You may be wondering "why just a tingle?" Well, when I first read about CMT coming to a server near me my geek johnson pitched a tent with enough head room for Yao Ming to enter without bending his head. But the truth of the matter is I am a cynic to the core. Vendors are notorious for contorting the truth and calling it marketing.  They all claim their product is the greatest thing since Internet porn which is obviously a lie. THERE IS NO THING BETTER THAN INTERNET PORN! There. I said it. Go tell your friends. I'll wait for you in the next paragraph....

So I'll read the technical specs and the research and whitepapers on a technology/product but I don't drink the marketing Kool-Aid. So until I get my hands on the box and start playing with it you'll only get a tingle out of me.

Friday, September 01, 2006

A Developer's Journal; Solaris 10 #1

I'm testing the Solaris 10 waters. My plan for this journal is to record my Solaris journey in as much useful detail as possible. But before I start reporting my trials and tribulations I want to discuss why I'm taking this path (remember Linux is my OS of choice).  There are 5 reasons why I'm looking at Solaris 10 now (from most to least importance):

  1. Try and Buy
  2. Chip Multi-Threading (CMT)
  3. DTrace
  4. ZFS
  5. Hotspot

Try and Buy

Try and Buy made the list and snatched the number one spot because it makes number 2 accessible given my limited R&D budget.

CMT

Sun has jumped ahead of the competition (Intel/AMD/IBM) in terms of throughput computing with the release of the T1 processor, code-named Niagara. I had previously discussed the coming of throughput computing here and here and I'm surprised yet thrilled to see it come to pass so rapidly. Most of the reviews of Niagara I've read are from people who don't have a clue about multi-threading and/or what makes it great. I'm guessing they have a rudimentary understanding of what a process is and probably don't know why scaling systems with threads is more efficient and performant than doing it with processes. So I'm really excited to get a crack at throwing an application that was designed from the ground up for multi-threaded deployment at Niagara.

DTrace

I'm a performance junkie. I like to make things go fast and it is hard to make something go fast when you don't understand what the thing is doing. Before DTrace it was possible to get pieces of the puzzle and sometimes it was possible to postulate about the whole and get it right but that's usually a hit or miss and it takes years of experience to develop the facilities to get more hits than misses. With DTrace there is no need to postulate it becomes possible to know with absolute certainty what the hell is going on. DTrace takes the guessing out of performance tuning.

ZFS

The disk subsystem is probably the most mission critical part of a computer because it's where all the programs and data are stored and without programs or data a computer isn't worth squat.

Aside:

Okay I concede. You can turn it into a modern art piece or make an ugly end table but other than that it isn't worth squat!

Unfortunately the disk subsystem is also the most unreliable. So any practical technology that handles reliability w/o major performance costs is a big deal to me. For years we've been subject to either uber expensive SCSI/RAID systems or cheap unreliable (and slower) IDE systems. But that's all changing, SATA and it's ilk are catching up to SCSI in terms of reliability and performance which will eventually drive SCSI out of the picture or force the price of SCSI down to the point where there is really no difference in terms of choosing one over the other. The bottom line is ZFS makes reliable disk subsystems something everyone can afford.

Hotspot

Given that the hotspot JVM is a Sun product and so is Solaris I conjecture that hotspot works better on Solaris than anywhere else. Especially Solaris on Sparc. So even though I'm a big Linux fan I'm an even bigger Java one.

Monday, June 12, 2006

QOTD: "I’m not writing shitty code; I’m creating refactoring opportunities."

I've got nothing to add to it. It's just plain funny as hell so I thought I would share it.

Saturday, June 03, 2006

"Look Ma, No Locks!"

Brian Goetz has written an excellent introductory article on nonblocking algorithms and showcases some simple nonblocking data structures with code examples and pictures.

The core concept you should take away from the article is the importance of CAS. It sits at the root of all nonblocking data structures in the Java platform.

Saturday, April 29, 2006

Amen!

Read the title.

Click on the link.

Read the post.

"That's all folks!"

Monday, April 24, 2006

Sublime

An interesting read to say the least.

Monday, April 10, 2006

New Template

I've just switched my blogger template so if you are a regular visitor don't freak out. The new template makes better use of screen real estate and is more code friendly.

Let me know what you think.

Saturday, April 08, 2006

regex vs. XML

I needed to screen-scrape wikipedia for the list of Top Level Domains for my email app so I could build an index of suffixes that would reduce the number of martian addresses the app would have to process. So I surfed over to the TLD page to have a look-see. Low and behold those beautiful wikipedians were nice enough to produce well formed XML (actually, XHTML 1.0 Transitional) content. My heart leaped w/ joy because no screen-scraping was needed. With a little bit of XPath I would have my index in no time.

WRONG!!!!!!!!!!!

I spent the next hour+ trying to get XOM to work and spit out my list via XPath because I sure as hell wasn't going to try to manually walk the tree extracting the TLDs. Do you have any idea how much code that would take? I had a feeling the reason it doesn't work has something to do w/ how XOM handles namespaces because when I access the tree the hard way it demands that I specify the namespace URI of the document to lookup any element in the tree even though the document declares one and only one namespace. Now this isn't a knock against XOM, though I'm sure if I had used dom4j it would have worked because dom4j is more concerned with being useful than being correct. But each have there place. By this time I was so aggravated I just didn't feel like downloading dom4j and unpacking it and setting my classpath, yada, yada, yada.

Regex to the rescue

This his how I got what I needed the old fashion screen-scraping way. 5 minutes tops from concept to results:
String tld = "http://en.wikipedia.org/wiki/List_of_Internet_top-level_domains"
Matcher regex = 
    Pattern.compile( "<td><a[^>]++>\\.(\\w++)</a></td>" ).matcher( "" );
BufferedReader reader =
    new BufferedReader(
        new InputStreamReader(
            new BufferedInputStream( new URL( tld ).openStream() ), "utf-8" ) );
PrintWriter writer = 
    new PrintWriter( 
        new BufferedWriter( 
            new FileWriter( new File( "/home/HashiDiKo/temp/TLD" ) ) ) );

for( String line; null != (line = reader.readLine()); )
    if( regex.reset( line ).matches() )
        writer.println( regex.group( 1 ) );

reader.close();
writer.close();

Saturday, April 01, 2006

This is truly cool

You have to check it out.