Monday, June 14, 2010

Idle = m + C1 + C2 + C3

There are a group of processor performance counters in Windows called "% Cx Time", which show the break down of processor idle time in which the processor goes into a low-power mode. Not all processors support these low-power idle states. The higher the number, the lower the power used by the processor, and the longer the latency before it can get back up running full speed again.

Thursday, June 10, 2010

CIL Constant Literals

Sometimes you'll want to create constant literals in CIL code. Most of the time you won't, but when you do, you'll need this. There are multiple ways of representing a constant, such as a constructor argument to a custom attribute, and ildasm.exe will give you something like this:
= ( DE AD BE EF )
If you want to re-write it, you know the signature (it's right next to the literal in the .il file,) so you only need to replace the single encoded binary literal into one like this:
= { string('Hello, World!') bool(true) type(int32) type([System.Xml]System.Xml.XmlDocument) }

Monday, June 07, 2010

Code Motion

The code we write in C# undergoes many changes before it is finally executed. There's the usual stuff interviews are made of: C# source is compiled to IL into a byte code representation that can be understood by the CLR. The CLR has another compiler - the JIT - which takes this byte code and turns each method into a sequence of instructions that can be understood by the hardware inside the computer. Then the hardware fetches the instructions from memory, decodes them, and executes them (fetching and writing to memory the program's data.) by driving voltages up and down across connected items inside your computer. If you thought your C# loop was executed as it's written, you're in for a treat: it's not.

The C# compiler produces different byte codes whether optimisation is turned on or off (off for DEBUG by default, which makes the emitted byte code look more like the source code, for easier pairing with the PDB files). The JIT compiler might take a loop and compile it down to a single inline sequence of instructions if they could make the execution faster. The processor might have multiple pipe lines in which it executes instruction streams in parallel (making use of the processor in times that it would otherwise have been wasting cycles waiting for the RAM or disk to send back data). Even after memory writes have been completed, they may be temporarily buffered close to the CPU rather than being expensively written directly to memory.

All of these changes are examples of code motion. To the average C# developer, most of these changes go unnoticed; on a single thread we have virtually no way of ever finding out that the instructions were actually executed out-of-order: the contract of C# guarantees these things for us. But in parallel processing with shared state, we start to witness evidence of the underlying chaos. Fields are updated after we expect - fields are updated before we expect. To limit any negative effects of code motion, the C# designers have given us the keyword, volatile. The keyword can be applied to instance/static fields (but not method local variables, as these cannot be safely accessed by multiple threads anyway - they will always live on their own thread's stack.) Read/write accesses to memory locations marked as volatile are protected by a "memory barrier". In oversimplified terms, any source code before a volatile read/write is guaranteed to have been completed before the access; similarly any source code after a volatile read/write will only be conducted after the access. To make your code ridiculously safe, you could add volatile to any updatable field (e.g. it can't be const or readonly) but you'd be losing out on all positive benefits of code motion.

Thursday, May 27, 2010

Dual Arduino Part 1 - Bus

I bought my first Arduino to help answer the question, "What can open source hardware teach me about running concurrent processes on my own computer?" The second was purchased to indulge my quest for knowledge of parallelism at a similarly low level, with Symmetric Multiprocessing (SMP). The third - a Mini - was just because I'm hooked. This is a post about the first and second. I'm only telling you about the Mini because of how great I think these little things are!

Threading: it's a technique to allow two logical processes to run concurrently, using shared physical computational resources. A program is usually logically separated into code and data. Code is executed, while data is read/written during execution (although executing data is not usually "done", it's completely possible). In a single-threaded program there is one call stack (representing the current method) and a bunch of special memory locations called registers, one of the registers "points" to the next instruction that will be executed. In a dual-threaded program, there must be two call stacks (one for each method currently being executed) and two sets of special memory locations (in one, the instruction pointer (IP) points to one part of the code; in the other it (usually) points somewhere different in the same code). However, in a single-processor system only one of these sets of memory locations can ever be stored in the physical registers at any given time, because there is only one set of physical registers. The other must be saved somewhere nearby, waiting in the wings. Something has to stop the current thread X, save its register values, restore Y's register values and start Y... e.g. the context switch.

SMP: two (or more) physical processors share the same bus (and thus access to all other hardware). With one thread, one processor would sit idle. With two threads, both processors get to do something (and unlike our single-processor example above, both sets of register values can now be physically stored in processor registers). With just two threads, we don't have to switch contexts and both processes can keep busy the whole time. Well, not all the time. While processor A is performing I/O operations against device C (reading/writing values in memory,) the bus route is busy and processor B must wait for access to C. Immediately, a memory bottleneck. Luckily, not all I/O needs to go over the bus every time: each processor has its own little local cache.

In the case of the Arduino, I don't have any fancy switching hardware to blur the distinction between buses and networks, so my bus is a set of parallel copper wires (some wires for addresses and some for data) connecting two Arduino devices to four EEPROM memory devices. Data written to an EEPROM by one processor can be read from the same EEPROM by another processor, and vice versa.

If both the Arduino devices are masters and all the EEPROM devices are slaves, how do we arbitrate access to the bus? How indeed. Something needs to send a clock signal. Something needs to stop both processors from trying to talk at the same time.

Monday, May 17, 2010

Reset IE Start Page

Here's a PowerShell snippet that will reset the home page in Internet Explorer; useful if you work somewhere where they always reset your home page to their own corporate one:
Set-ItemProperty
-path "HKCU:\Software\Microsoft\Internet Explorer\Main"
-name "Start Page"
-value "http://www.google.co.uk"

Tuesday, May 04, 2010

Yet Another Post About Compression in IIS 6.0

By default, IIS won't try and compress your dynamically generated ASPX files even if you've got dynamic compression enabled as per my earlier posts. Go ahead and include the ASPX extension like this:
cscript C:\Inetpub\AdminScripts\adsutil.vbs SET W3SVC/Filters/Compression/Deflate/HcFileExtensions htm html txt css js
cscript C:\Inetpub\AdminScripts\adsutil.vbs SET W3SVC/Filters/Compression/Deflate/HcScriptFileExtensions asp dll exe aspx

Sunday, May 02, 2010

Parallelism

It's not a word that rolls of your tongue. Try explaining - without appearing to have recently suffered a blow to the head - to a room full of people that turning down your SQL Server's maximum degree of parallelism may solve some of their performance woes.

Parallelism (this definition is stolen from the beginning of Joe Duffy's book, Concurrent Programming on Windows) is the use of concurrency to decompose an operation into finer grained constituent parts so that the independent parts can run [separately].

Parallelism is intended to take advantange of physical resources that might be otherwise unused. If you have a large table or index that needs to be scanned, and you have 24 available cores, SQL Server might want to parallelize your query into 24 concurrent tasks. In this case, it usually leads to resource starvation, how many database servers are sitting idle all the time, waiting for your "embarrassingly parallel" query to arrive. Do you need to put a question mark at the end of a rhetorical question. On a SQL Server, Microsoft's official word is to keep the maximum degree of parallelism (dop) below 8, or the number of cores in a single socket, whichever is lower. Don't forget that in the example above, your server would need 24x the RAM required to turn the sequential operation into a concurrent one.

PFX is a .NET library for enabling parallelism in your C# code. Although I haven't got to the bottom of it, I'm tempted to believe that it - too - would depend entirely on the load profile of your server, so give it a try and make sure you test it with an adequate cross section of concurrent users to make sure that the parallelism selfishness isn't the cause of any new performance issues.

Monday, April 26, 2010

Hello, Axum!

using System;
using Microsoft.Axum;

public domain HelloWorld {
public agent Program : channel Application {
public Program() {
string[] args = receive(PrimaryChannel::CommandLine);
Console.WriteLine("Hello, World!");
PrimaryChannel::Done <-- Signal.Value;
}
}
}

Saturday, April 10, 2010

Jerk: Scalability and Higher Order Derivatives

I've long thought of scalability as the first derivative in the relationship between resources and operations per unit of time. With no resources, we usually achieve no operations, regardless of how long. With a single resource, we might achieve Om operations over a time span of T. By adding a second resource, we might achieve On operations over the same time span. We would say that the marginal scalabity between Rm (1) and Rn (2) was (On - Om) / (Rn - Rm). Looking at Operations vs. Resources on a line chart, this marginal scalability would be the slope of the line connectiong ORm and ORn. Knowing your application's scalability between m and n is a starting point, but the probability of ideal linear scalability in a real world application is 1/infinity. We need to know what happens when you add a third resource. This (second derivative) will tell us the rate at which the scalability changes as more resources are applied, and can be a much more useful figure. It may take a little more time until I can clarify what jerk might meaningfully contribute to discussion of scalability, but essentially it's the change in acceleration. :-)

Wednesday, April 07, 2010

WebAdministration

If you're running Windows 7, PowerShell 2.0 and IIS 7.5 you may wonder how to get started with modifying your IIS configuration using PowerShell scripts (the iis.net web site is a bit unclear, the web installer gives an unhelpful lie phrased as an error message, and the downloadable install gives up without being much help either.)
No need for extra installs beyond the usuall Windows Features -> Internet Information Services install, just follow these steps:

  • "Run as Administrator" when starting PowerShell.
  • Set-ExecutionPolicy RemoteSigned (or whatever blows your hair back)
  • Install-Module WebAdministration
  • Get-PSDrive
  • cd IIS:

Saturday, April 03, 2010

F#ing Around

In my last post, I wrote a tiny standalone F# program that would print out a reversed string. It wasn't enough, I wanted to call the function from C#. Difficult? Not at all, but there are a couple of tricks that might catch you at first.
1) Add a new C# project to the solution
2) Add a new assembly reference to the F# project
3) One of two (or is that three?) things:
a) did you name your F# module (a declaration at the top of the .fs file like module MyFirstGreeting? If yes, you will reference the function from C# by it's fully qualified name: MyFirstGreeting.reverse(...)
b) do you have an unnamed F# module? If yes, it's assumed the name of the .fs file, most likely Program, so you'll need to invoke the function by calling Program.reverse(...)
c) if you answered yes to b) and the class in your C# project was also called Program, you are waiting for this: use the global namespace qualifier! global::Program.reverse(...). Remember your C# Program's fully qualified class name is actually something like ConsoleApplication2.Program.

Compile and run! Happy days! And marvel in the wonder of Visual Studio 2010 as you step through the code from C# to F# (the call stack window even shows what language your current stack frame is written in!). This is going to be a fun journey...

Monday, March 29, 2010

Hello F#

I just got my Release Candidate of Visual Studio 2010 up and running so I thought I'd do the obligatory recursive greeting function (not as pretty as Ruby or Python, but we're talking about F# here!)
let rec reverse (value : string) =
if value.Length < 2 then
value
else
value.Chars(value.Length - 1).ToString() + reverse(value.Substring(0, value.Length - 1))

[<EntryPoint>]
let main (args) =
printfn "%s" (reverse("!dlroW ,olleH"))
0

Tuesday, March 23, 2010

Installing memcached

On ubuntu:
* download the memcached source from memcached.org (it's on Google Code at http://code.google.com/p/memcached/)
* download the libevent source from http://www.monkey.org/~provos/libevent/
* open a terminal window
* cd to where the two tar.gz files are
* tar -zxvf memcached-1.4.4.tar.gz
* tar -zxvf libevent-1.4.10-stable.tar.gz
* cd into the libevent folder
* ./configure && make
* sudo make install
* cd up and into the memcached folder
* ./configure && make
* sudo make install
* at this point we could try starting memcached, but it would probably fall over with an error about . we need to give the memcached an easy way to find libevent. (running whereis libevent is good enough for us to verify where it's actually been installed: I suspect /usr/local/lib)
* sudo gedit /etc/ld.so.conf.d/libevent-i386.conf
* on it's own line in the file, enter the following, and save and quit gedit:
* /usr/local/lib/
* even though you may have just installed it on an amd64 box, the file name needs to be i386
* sudo ldconfig is the final set up step
* now fire up memcached in verbose mode (use the -vv parameter)
* alternatively run it in daemon mode (as it was intended, with the -d parameter)
* once it's up and run you can muck around with it by starting a telnet session
* telnet localhost 11211
* try some of the commands in the very helpful reference http://lzone.de/articles/memcached.htm
* :-)

Monday, March 22, 2010

"Velocity" Part 1

It's just the CTP3 version, but it's nearly rough enough to put me off new technology for a while. I eventually got it set up on a single machine using SQL Server for its configuration mechanism, which necessitated creating a login, a database named "VelocityCacheConfig" and linking the login with a user "Velocity" who was db_datareader, db_datawriter and db_owner.

Pre-requisites: .NET 3.5 SP1 and Windows PowerShell.

You need to run the supplied PowerShell script as admin. Fair's fair. It's intended usage is to manage the cache cluster and you wouldn't want just anybody doing that.

There was some error message about the firewall.

It created a new region for every item added to the cache.

It took a long time to Remove-Cache and you can't New-Cache with the same cache name until the old one's removed. Get-CacheHelp lists the PowerShell cmdlets for working with the cache.

Need to add a reference to ClientLibrary.dll and CacheBaseLibrary.dll to your Visual Studio project.

Before running the following C#, you'd have to Start-CacheCluster, and Add-Cache "Test".

using Microsoft.Data.Caching;

DataCacheServerEndpoint[] endpoints = new DataCacheServerEndpoint[]
{
new DataCacheServerEndpoint("SCANDIUM", 22233, "DistributedCacheService")
};
DataCacheFactory factory = new DataCacheFactory(endpoints, true, false);
DataCache cache = factory.GetCache("Test");

Wednesday, March 10, 2010

SqlBulkCopy

SQL Server 2008 has table-valued parameters in ADO.NET, but if you're stuck with a poorly performing inserts and a previous version of SQL Server, you can still get some awesome speed with the System.Data.SqlClient.SqlBulkCopy class. Just create a new one (in a using block), set up the destination table name and optionally a set of column mappings, and fire away. I ran a test with 65536 rows (width: 3 x INT + 1 x FLOAT) and it completed in 750ms compared with 36'447ms to send the equivalent table row-by-bleeding-row using multiple stored procedure invocations. Also, it's much easier to use than writing your own wrapper around bcp.exe (the only option if you're stuck with Sybase)!

Tuesday, March 09, 2010

Unicode

UTF-16 is a way of representing all of the UCS code points in two bytes (or four). It can encode all of the code points from the Basic Multilingual Plane (BMP) in just two bytes, but code points in other planes are encoded into surrogate pairs. UCS-2, as used by SQL Server for all Unicode text, was the precursor to UTF-16 and can only handle code points from the BMP. It is forward compatible with UTF-16, but any code point outside of the BMP encoded in UTF-16 will appear to be two separate code points *inside* the BMP if the encoding is UCS-2. The data will be preserved - it is only the semantics (i.e. the abstract "character(s)") that differ.

Note: This is extremely unlikely for modern business applications, as the languages outside the BMP are academic, and/or historical, such as Phoenician. For these purposes, UCS-2 and UTF-16 can be considered equivalent and interchangeable.

Side note: UTF-8 is just another way of representing all of the UCS code points from all of the planes, but the bits are encoded in a different way to UTF-16.

Extra note: .NET uses UTF-16 for its in memory encoding of the System.String type. A .NET System.Char is limited to 16-bits, and therefore a single char cannot hold a UTF-16 encoded surrogate pair (e.g. any code point not on the BMP). The char data type is similar to UCS-2 in this respect. Getting the char at a specific index of a string will always return a single 16-bit char, essentially breaking up a surrogate pair if one existed at this index of the string.

Saturday, March 06, 2010

Alternative to Clustering

Today my desktop machine had a hardware failure. Everything froze up, I tried rebooting and all I got was a symphony of beeps from the BIOS. Nada. If this had been a production database server and it wasn't in a cluster (as you can tell I'm having a lot of clustering problems lately) I would have been up the creek without a paddle. As it happens, I have a spare box lying around with very similar spec, and I was able to take my boot disk out of box 1 and put it into box 2 with very little downtime. It got me thinking. Even if you don't have a cluster with a warm machine waiting for failover, you could still reduce the cost of the outage by keeping a cold server of similar spec waiting for just such an occasion. If the data and log files are stored in a SAN, you could theoretically just bring up the cold server and attach the databases, and do some DNS magic on the clients... Would it work? Well, isn't that what DR days are for?

System.IO.Compression and LeaveOpen

I got into the habit of stacking up blocks of using statements to avoid lots of nesting and indentation in my code, ultimately making it more readable to humans!

using (MemoryStream stream = new MemoryStream())
using (DeflateStream deflater = new DeflateStream(stream, CompressionMode.Compress))
using (StreamWriter writer = new StreamWriter(deflater))
{
// write strings into a compressed memory stream
}


The above example shows where this doesn't work though. The intention was to compress string data into an in-memory buffer, but it wasn't working as expected. There was a bug in my code!

When you're using a DeflateStream or a GZipStream, you're (hopefully) writing more bytes into it than it's writing to its underlying stream. You may choose to Flush() it but both streams are still open and you can continue to write data into the compression stream. Until the compression stream is closed, however, the underlying stream will not contain completely valid and readable data. When you Close() it, it writes out the final bytes that make the stream valid. By default, the behaviour of the compression stream is to Close() the underlying stream when it's closed, but when that underlying stream is a MemoryStream this leaves you with a valid compressed byte stream somewhere in memory that's inaccessible!

What you need to do instead is leave the underlying stream open, using the extra constructor argument on the compression stream, like this:

using (MemoryStream stream = new MemoryStream())
{
using (DeflateStream deflater = new DeflateStream(stream, CompressionMode.Compress, true))
using (StreamWriter writer = new StreamWriter(deflater))
{
// write strings into a compressed memory stream
}
// access the still-open memory stream's buffer
}

Friday, March 05, 2010

Powershell Script for New Guid

> function NewID { [System.Console]::WriteLine([System.Guid]::NewGuid()); }
> NewID
83ced965-68e6-454d-aa8e-2f056ae1a030

Money and Trouble

Use "Swim Lanes" to isolate the money makers and to isolate the trouble makers. The principle is intended to give higher quality service to paying customers, or more specifically the ones on whom your business is more dependent. It is also intended to isolate failures in troublesome components so that they do not affect the overall system negatively.