Search This Blog
19 February 2010
cfgadm
Checking 'zpool status' I was able to verify that one of the disks had died. I'm running raid-z2, so it isn't like it was a hectic problem to solve... but I learned long ago it is better to fix it now than to wait until a couple more die.
Normally, you would have to go buy a replacement disk. Luckily for me, this Asus box never did recognize the last 3 drive bays. Counting across (there are 10 hot-swap SATAII drives -- want to make sure to pop the right one) I took a guess which one was failing and which ones were not currently in use. Luckily I was right.
Now, the "new" drive already had data on it because it used to be part of the root mirror before I upgraded them to larger drives. That being said, I was a little confused when the new drive wasn't being recognized.
I found this page which helped dramatically. 'cfgadm' showed that drive '6/0' was not configured. I ran 'cfgadm -c configure sata6/0'. It now showed up, but it said it was 'unavail' and 'corrupted data'. 'zpool online' didn't work because of those errors. Finally, I managed to get it to start working with 'zpool replace -f data c6t0d0'. It took quite awhile for it to finally finish. 'fmadm faulty' still showed the fault but I was able to fix that with the zpool clear that it recommended.
I'm thinking I should hook up one of my woot-off lights to flash whenever fmadm faulty shows a failure...
22 July 2009
Western Digital My Book World
I tried rebooting it, power cycling it, etc. No luck.
I contacted support. So far there has been 6 responses from Tech Support (which of course) started off with:
My apologies; I failed to notice you are using Linux.Western Digital technical support only provides jumper configuration and physical installation support for hard drives used in systems running the Linux/Unix operating systems. For setup questions beyond physical installation of your Western Digital hard drive, please contact the vendor of your Linux/Unix operating system.Which of course is *really* helpful since it is a *Linux* NETWORK drive accessed via web browser and samba. Of course, I understand they aren't trained to understand that a browser or ping on a non-windows system is just as accurate as on a windows system, but...
Still no resolution. I am expecting they are going to tell me to RMA the drive -- but I have a lot of personal sentimental pictures and such on there that have no backups (as these were the backups to the Windows box that died)... No clue what I am going to do -- but since I seem to have to RMA every Western Digital device I buy (internal or external) maybe I should quit buying WD.
07 July 2009
Hanging during boot
It still takes 4-5 hours to boot.... specifically, dmesg shows that this took 4 hours...
Jul 7 08:15:32 serveris xpv_psm: [ID 805372 kern.info] xVM_psm: ide (ata) instance 1 irq 0xf vector 0x98 ioapic 0x4 intin 0xf is bound to cpu 3
Jul 7 08:15:32 serveris xpv_psm: [ID 805372 kern.info] xVM_psm: ide (ata) instance 1 irq 0xf vector 0x98 ioapic 0x4 intin 0xf is bound to cpu 0
Jul 7 08:15:40 serveris xpv_psm: [ID 805372 kern.info] xVM_psm: ide (ata) instance 1 irq 0xf vector 0x98 ioapic 0x4 intin 0xf is bound to cpu 3
Jul 7 08:15:57 serveris xpv_psm: [ID 805372 kern.info] xVM_psm: ide (ata) instance 1 irq 0xf vector 0x98 ioapic 0x4 intin 0xf is bound to cpu 0
Jul 7 08:15:57 serveris xpv_psm: [ID 805372 kern.info] xVM_psm: ide (ata) instance 1 irq 0xf vector 0x98 ioapic 0x4 intin 0xf is bound to cpu 1
Jul 7 08:15:57 serveris xpv_psm: [ID 805372 kern.info] xVM_psm: ide (ata) instance 1 irq 0xf vector 0x98 ioapic 0x4 intin 0xf is bound to cpu 2
Jul 7 08:15:57 serveris xpv_psm: [ID 805372 kern.info] xVM_psm: ide (ata) instance 1 irq 0xf vector 0x98 ioapic 0x4 intin 0xf is bound to cpu 3
(repeat and repeat and repeat)
Jul 7 12:21:14 serveris xpv_psm: [ID 805372 kern.info] xVM_psm: ide (ata) instance 1 irq 0xf vector 0x98 ioapic 0x4 intin 0xf is bound to cpu 1
(system becomes usable here)
Jul 7 15:24:19 serveris xpv_psm: [ID 805372 kern.info] xVM_psm: ide (ata) instance 1 irq 0xf vector 0x98 ioapic 0x4 intin 0xf is bound to cpu 3
Any thoughts why a non-existant IDE device is polled for 4 hours during boot? Or why after finally booting, it is *still* doing it?
UPDATE: Bug submitted
20 May 2009
OpenSolaris Bible: First Impressions
First thing I decided to play with was getting those remote CIFS shares mounted (they had been accessible via file explorer, but not via command-line). Unfortunately, it required I reboot the machine after enabling the smb/server... and well, for some unknown reason that currently takes 4 hours... so I didn't get to play with it much. I did however manage to get one of the shares to mount... Still trying to get the automount to work, but it's a start.
Taking the book to work today, I decided to play with FMD... Here's some comparison results between the work machine and the home machine... [Note: 'eris' is the work machine. 'serveris' is the home machine]
[eris:malachi(0)] ~% pfexec /usr/lib/fm/fmd/fmtopo
TIME UUID
May 20 09:40:20 23d4dce1-d973-c98a-ee18-ee2a16dd8ee8
hc://:product-id=OptiPlex-755:chassis-id=72VYGH1:server-id=eris/motherboard=0
hc://:product-id=OptiPlex-755:chassis-id=72VYGH1:server-id=eris/motherboard=0/chip=0
hc://:product-id=OptiPlex-755:chassis-id=72VYGH1:server-id=eris/motherboard=0/chip=0/core=0
hc://:product-id=OptiPlex-755:chassis-id=72VYGH1:server-id=eris/motherboard=0/chip=0/core=0/strand=0
hc://:product-id=OptiPlex-755:chassis-id=72VYGH1:server-id=eris/motherboard=0/chip=0/core=1
hc://:product-id=OptiPlex-755:chassis-id=72VYGH1:server-id=eris/motherboard=0/chip=0/core=1/strand=0
hc://:product-id=OptiPlex-755:chassis-id=72VYGH1:server-id=eris/motherboard=0/hostbridge=0
hc://:product-id=OptiPlex-755:chassis-id=72VYGH1:server-id=eris/motherboard=0/hostbridge=0/pciexrc=0
hc://:product-id=OptiPlex-755:chassis-id=72VYGH1:server-id=eris/motherboard=0/hostbridge=0/pciexrc=0/pciexbus=1
hc://:product-id=OptiPlex-755:chassis-id=72VYGH1:server-id=eris/motherboard=0/hostbridge=0/pciexrc=0/pciexbus=1/pciexdev=0
hc://:product-id=OptiPlex-755:chassis-id=72VYGH1:server-id=eris/motherboard=0/hostbridge=0/pciexrc=0/pciexbus=1/pciexdev=0/pciexfn=0
hc://:product-id=OptiPlex-755:chassis-id=72VYGH1:server-id=eris/motherboard=0/hostbridge=1
hc://:product-id=OptiPlex-755:chassis-id=72VYGH1:server-id=eris/motherboard=0/hostbridge=1/pciexrc=1
hc://:product-id=OptiPlex-755:chassis-id=72VYGH1:server-id=eris/motherboard=0/hostbridge=2
hc://:product-id=OptiPlex-755:chassis-id=72VYGH1:server-id=eris/chassis=0
malachi@serveris[0]:~ % pfexec /usr/lib/fm/fmd/fmtopo
TIME UUID
May 20 09:42:42 0524f912-299c-6280-d892-c36f53c00103
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/chip=0
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/chip=0/core=0
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/chip=0/core=0/strand=0
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/chip=0/core=1
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/chip=0/core=1/strand=0
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/chip=0/memory-controller=0
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/chip=0/memory-controller=0/dram-channel=0
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/chip=0/memory-controller=0/dram-channel=1
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/chip=1
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/chip=1/core=0
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/chip=1/core=0/strand=0
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/chip=1/core=1
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/chip=1/core=1/strand=0
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/chip=1/memory-controller=0
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/chip=1/memory-controller=0/dram-channel=0
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/chip=1/memory-controller=0/dram-channel=1
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=0
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=0/pcibus=5
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=0/pcibus=5/pcidev=11
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=0/pcibus=5/pcidev=11/pcifn=0
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=0/pcibus=5/pcidev=32
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=0/pcibus=5/pcidev=32/pcifn=0
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=1
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=1/pciexrc=1
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=2
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=2/pciexrc=2
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=3
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=3/pciexrc=3
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=3/pciexrc=3/pciexbus=2
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=3/pciexrc=3/pciexbus=2/pciexdev=0
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=3/pciexrc=3/pciexbus=2/pciexdev=0/pciexfn=0
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=4
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=4/pciexrc=4
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=4/pciexrc=4/pciexbus=1
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=4/pciexrc=4/pciexbus=1/pciexdev=0
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=4/pciexrc=4/pciexbus=1/pciexdev=0/pciexfn=0
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=5
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=5/pciexrc=5
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=6
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/hostbridge=6/pciexrc=6
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/chassis=0
first cpu:
[eris:malachi(0)] ~% pfexec /usr/lib/fm/fmd/fmtopo -P all hc://:product-id=OptiPlex-755:chassis-id=72VYGH1:server-id=eris/motherboard=0/chip=0
TIME UUID
May 20 09:44:10 2d75b8fe-0c11-4e89-d8a5-ed2569e7b58e
hc://:product-id=OptiPlex-755:chassis-id=72VYGH1:server-id=eris/motherboard=0/chip=0
group: protocol version: 1 stability: Private/Private
resource fmri hc://:product-id=OptiPlex-755:chassis-id=72VY...
group: authority version: 1 stability: Private/Private
product-id string OptiPlex-755
chassis-id string 72VYGH1
server-id string eris
group: chip-properties version: 1 stability: Private/Private
vendor_id string GenuineIntel
family int32 6
model int32 23
stepping int32 6
malachi@serveris[0]:~ % pfexec /usr/lib/fm/fmd/fmtopo -P all hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/chip=0
TIME UUID
May 20 09:45:56 a0c248a0-75ae-65c3-e8b3-be99702ad460
hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-1234567890:server-id=serveris/motherboard=0/chip=0
group: protocol version: 1 stability: Private/Private
resource fmri hc://:product-id=L1N64-SLI-WS:chassis-id=SYS-...
group: authority version: 1 stability: Private/Private
product-id string L1N64-SLI-WS
chassis-id string SYS-1234567890
server-id string serveris
group: chip-properties version: 1 stability: Private/Private
vendor_id string AuthenticAMD
family int32 15
model int32 193
stepping int32 3
NodeId uint32 0x0
CoherentNodes uint32 0x2
SbNode uint32 0x0
LkNode uint32 0x0
SystemCoreCount uint32 0x4
C0Unit uint32 0x0
C1Unit uint32 0x1
McUnit uint32 0x2
HbUnit uint32 0x3
SbLink uint32 0x1
BroadcastRoutes uint32[] [ 9 1 ]
ResponseRoutes uint32[] [ 1 8 ]
RequestRoutes uint32[] [ 1 8 ]
shorthand for the CPU names....
[eris:malachi(0)] ~% pfexec /usr/lib/fm/fmd/fmtopo -s cpu
TIME UUID
May 20 09:45:57 08533a3a-65de-e73d-e8f1-e89fba223ed0
cpu:///cpuid=0
cpu:///cpuid=1
malachi@serveris[0]:~ % pfexec /usr/lib/fm/fmd/fmtopo -s cpu
TIME UUID
May 20 09:47:22 d2a17d23-222d-412d-c691-cf1273371317
cpu:///cpuid=0
cpu:///cpuid=1
cpu:///cpuid=2
cpu:///cpuid=3
Now, what else can I play with?
10 May 2009
Server Died

So I got home last night to find the server had power but was not responding (mouse, keyboard, ssh, anything). I tried rebooting, but it kept handing at "Hostname: serveris" and wouldn't go any further (even in single-user mode). I saw some chatlogs online that suggested adding '-k -a -d verbose' and using '/dev/null' to the answer of any questions (like /etc/system replacement)...
I tried that and got this far (see left)...
Looking around some more, I saw that if I changed the '-k' to '-kd' it would drop it into debug mode. At that point, I did the following:

[0] moddebug/W 80000000
[0] :c
This allowed me to see a few more details.... (sorry for the blurriness of the pic - it was about 3am)
After trying to find anything online that would help (and the IRC channel) I finally said screw it and decided I would reinstall opensolaris on the root mirror.
I downloaded the USB version of OpenSolaris 0906 111a, but evidentially my quad core machine does not have the option of booting from USB (WTF?). I reburned the CD version and installed it. One thing that confused me is that although my old system was 10/08 upgraded to 111a and the new version was supposed to be 0906 111a, it now says 101b.
Trying to boot the new one, it again hung. At a different position, but... I was starting to think it was a hardware problem. I let it try to boot overnight and the next morning it was finally at the login prompt... with the old install.
The logs showed that it had tried to load the Belkin UPS a few steps after where it locked up, so I unplugged the UPS. I went ahead and applied all updates and rebooted. It took about 5 hours for it to finally boot again (though it did). It still says I am using 101b and that there are no new updates.
The xVM instance is there and I was able to start it. The whole root zone however is gone. The ZFS partition is there, and empty. zoneadm doesn't show anything but global. So, I am going to try to recreate the global zone, but... I still don't know what happened. I am also concerned that it currently takes about 5 hours to boot.
13 November 2008
Printers won't Print
[ Nov 13 13:13:10 Executing start method ("/lib/svc/method/svc-network-discovery start snmp"). ]
/usr/bin/dbus-send --system --print-reply --dest=org.freedesktop.Hal --type=method_call /org/freedesktop/Hal/devices/network_attached org.freedesktop.Hal.Device.NetworkDiscovery.EnablePrinterScanningViaSNMP int32:60 string:public string:0.0.0.0
Error org.freedesktop.DBus.Error.NoReply: Did not receive a reply. Possible causes include: the remote application did not send a reply, the message bus security policy blocked the reply, the reply timeout expired, or the network connection was broken.
After searching around for awhile, I found that this bug had already been reported.
So first I installed 'SUNWsmmgr' via the package manager
root@eris:~# svcadm restart svc:/system/hal:default
root@eris:~# svcadm clear svc:/network/device-discovery/printers:snmp
root@eris:~# svcadm enable svc:/network/device-discovery/printers:snmp
That seemed to fix it.
As far as setting the printers up, it seemed to work much quicker setting it up using socket, the IP, default port (as opposed to my last attempt using samba and the name)....
28 August 2008
Memory Upgrade
24 March 2007
Server Built

Brett came by to help me get the server built yesterday. I have been waiting for Solaris 10 b60 (which has AMD-V support), but I am starting to seriously consider using b59 to install and then upgrading.
Only real problem with the build was that the Armor Extreme case had drive rails that got in the way of the Addonics drive cages... Brett managed to dremmel them out of the way.
On the left you will see 4 pictures... Yes, I resized it (Brett hates those 13MB photos), but you can still click on it to get a slightly larger photo.
The left images are with the doors open. The right images are with the doors closed.
The top images are with a flash (see detail) and the bottom without (see lights)...
Now to choose an installation medium :)
21 February 2007
Belkin F1DE208C
Unfortunately, it doesn't appear to work. Both 2x7-segment LEDs say "88" and the rest of the lights are randomly flashing...
I would look up the error code, but the manual has nothing, the website has nothing, and I am not finding anything on Google. Guess I will have to talk with Belkin support tomorrow.
Good thing I have the laptop.
16 February 2007
05 February 2007
New Hardware Purchased
Asus L1N64-SLI WS
2x AMD FX-70
Thermaltake VA8004BWS
Thermaltake HardCano 13
2x Kingston KVR667D2E5K2/2G
Thermaltake 850Watt W0131
going to use an existing space nVidia 6600GT
2x Hitachi Deskstar T7K500 250GB Serial ATA II (mirrored boot)
5x Hitachi Deskstar T7K500 250GB Serial ATA II (raidz2 data)
2x Addonics AE5RCS35NSA
NEC Black Floppy (memtest86, hitachi drive tool, etc etc)
going to use an existing dual layer burner
Belkin 1500VA/1000 joules UPS
We're also upgrading the KVM:
Belkin F1DE208C
3x Belkin F1D9400-06 (for machines on the rack)
1x Belkin F1D9400-25 (for docking station on the desk)
Still need to grab a Cat6 cable from PCHCables when we get a chance to run out to Hillsboro.
Plan is to run OpenSolaris as Xen dom0, then host all Xen domU via locally shared ZFS NFS directory...
01 February 2007
30 January 2007
Plan shot down - looking at new plan
Which is wierd, since cases are usually around $20 and the drives (250GB Samsung) were only $52 and the OS (NetBSD or Linux) is free...
I was looking at possibly making my own via mini-itx or something. Found iSCSI-SCSI bridges (only $1200)...
maybe it is time to make my own...? well, I would, but I really don't have the time to experiment right now... so I guess I once again have to get a larger server box up and running; then take my time designing my iSCSI solution (rotozip must like the enclosures)... And THEN design my nice small cool (as in temp) server....
In the meantime, got to figure out what to do... build new AMD-V machine (which isn't on Solaris HCL) or convert the last machine I bought (that has memory controller bugs and Asus lies all over it)
29 January 2007
25 January 2007
Possible Plan
Then the Solaris dom0 sets these up as Raid-Z.
Solaris then exports a NFS root for each domain. In fact, we could also export CDROM images and such over NFS so that the various OSes could have access to them.
At this point we would either have a small grub floppy image that would do network boot; or we have Xen boot the domU directly over the network. In either case, the plan is the same, have the OS run 'diskless' over NFS (with rw permissions)... If they have access to change which kernel grub points to, probably better.... so what if we have a small local grub image that points to a remote NFS grub image that is local to that filesystem... and thus they could change which kernel image, but not where it is...
Anyways, then, if we want to increase storage capacity, we add a drive to the network (iSCSI) and tell ZFS to add it to the pool... voila, all the domains have more space.
And if we want to back up a domain, we can take a snapshot and clone, etc.... I see us taking a snapshot at say 12:01am on monday, then incremental snapshot every day; thus we can rollback up to a week... although, probably need to maintain 2 weeks if we want to restore 1 -- otherwise every monday when we kill the primary full snapshot, all the incrementals become useless...
And if there is a hardware failure, Raid-Z should autorecover and provide means for me to fix it.
And if we need a test domain, we should be able to snapshot/clone (although booting with the same network parameters could be a problem).
And we should be able to write a tool (web based) that they can run in their own domain to request that we revert back to a previous snapshot...
And if we need to, we can specify quotas or reservations on a per domain or even on a per directory...
And although NFS should cause some overhead, since it is on the same physical server, I doubt it will even talk to the network...
And if we ever need to offload one of the domU to a different server, it should be extremely simple since we'd only need to move the Xen configuration and the tiny grub floppy image.
Overall I am liking this idea. Not sure how well it will work, but I guess I should start looking at hardware for it.
BTW: The Solaris dom0 needs redundancy too. I figure that what we would do is have a small mirrored root to boot Solaris -- but the dom0 needs to be the one in charge of the Raid-Z, so...
23 January 2007
Thoughts... ie: What should I do?
* We do not want a repeat of the data wipeout we got a year ago... As such, if one domain gets hacked, it can not take down the rest of the domains.
- I had been planning on using Jail to alleviate this problem, but at this point I think our best bet is Xen
- While many OSs support Xen domU, many fewer support Xen dom0.
* Also, if a domain gets hacked and wiped out, we need a backup to restore from.
- ZFS Snapshots (maybe full weekly and incremental daily) seems like a good solution for this
- SVN might be a good solution to this
- It might be good if there was a way to automate this (ie: allow the root user to do it instead of emailing me)
* We want to get disk redundancy, while keeping cost and maintenance low.
- RAID-0 gives us the redundancy, but at a fairly high cost. The current 1TB gives us less than 500GB this way.
- RAID-5 could be acceptable, as long as we could fix the system if the root partition were the one to screw up
- ZFS and Raid-Z seem like excellent options to provide the redundancy and automatic corrections
* Each domain should be able to have it's own root user and OS.
- Jail could have provided for this, but seems very buggy ('man man' in a jail crashed the entire OS)
- Xen should allow for this fairly easy. As such, the base OS really doesn't do anything other than launch the domUs.
* If multiple OSs share a base OS, it would be nice (though not required) if we were able to not waste a lot of duplicated space
- Unionfs is ideal for this. Not sure how easy it is to use with Xen
- At some point, it has to be better to have a full copy (like if they upgraded every single app)
* It would be preferable if all OS's could run unmodified, thus allowing us greater flexibility and choices
- Intel VT or AMD-V technology. If we are planning on doing any RSA work on any domain, it can't be the Intel option. So probably an Opteron 2xxx (or 2).
* If we run low on drive space, it should be fairly simple to add more
- ZFS is supposed to allow for this with 'zfs add'
- We *could* do separate drives for each domain, but that would waste a lot of drive space
- iSCSI might be a better option than internal drives for this. Instead of running out of hard drive bays, we'd just have to worry about running out of ethernet ports -- which realistically, is much easier to come up with more of. ** SEE BELOW
** iSCSI Thoughts:
- IDEALLY, I would have a set of cheap 250GB or 500GB drives that host themselves as iSCSI Targets. Then, adding new drives is literally adding another drive to the network.
- We could, in theory, run Solaris or NetBSD to easily provide iSCSI Targets to the network - but then we have to wonder why we are creating a separate box to host Xen
- If we have multiple iSCSI Targets, which our Solaris(?) box were to RAID-Z together, we could also have extra iSCSI Targets that VMWare and/or other hardware could use directly (even Windows)
- If we have multiple iSCSI Targets, which our Solaris(?) box were to RAID-Z together, the Solaris box could provide the pool as a iSCSCI Target to other devices/OSes
- If we were to have Solaris running RAID-Z directly on the iSCSI Target (ie: providing ZFS pools as iSCSI Targets), then separate Xen domUs could point to these iSCSI Targets for their primary data storage and POTENTIALLY get the ZFS benefits regardless of their OS (snapshots, error-correction [at least from hardware errors], etc)
- In theory, if we could come up with mini embeddable systems that could boot from iSCSI Targets, then we could do away with Xen entirely and just have a little device on the network for each domain... this might be overkill
- Note: Can Solaris boot off iSCSI?
22 January 2007
Another thing to consider
16 January 2007
EMail conversation with Sun
Hi Malachi,My response:
Thank you for visiting the web chat forum on the Sun website yesterday, when you expressed interest in the Try and Buy promotion being currently run by Sun.
You mentioned that you are deciding which machine you will trial on the promotion - excellent! You will see the full list of products on the promotion at www.sun.com/tryandbuy
You can email me with any questions that you may have, or you can visit our team at the web chat forum. Here's some additional information on the Try and Buy promotion, which you may find useful:
When you receive the Server on the Try and Buy promotion you will also receive a Welcome Pack.
Inside the Welcome Pack you will find-
1. A Quick Guide to Installation.
2. Access to Tuning and Optimization documentation.
3. Access to an online portal and forum for system and application tuning hints and tips.
4. Access to Sun's performance engineers.
5. All of the latest patches and configuration files for the system on the portal.
Additional options for purchase are available for system ready and business ready services.
The Solaris 10 Operating System and other software provided with the system is covered under the Warranty support during the Try and Buy period.
Best Regards
Mark Cradock
Sun Microsystems
Tel: + 353 (0)599136768
mark.cradock@your-sun.com
www.sun.com
Hi Mark,
I am debating the Sun Fire X4200 or the Sun Ultra 40 Workstation. Perhaps it would be better for me to explain the intended usage and get your feedback.
First, a little background. We had a FreeBSD box running for about 4 years without reboot. As is probably obvious from that statement, we were behind on the security updates. While we were hosting multiple domains, they were not in Xen or Jail or anything. On Christmas eve a year ago, a hacker from a Polish cable isp hacked in through one of the user accounts and wiped out the entire system. Because of this, I decided this time that I wanted to ensure that the various domains stay secure even if one is compromised.
We bought the latest top-of-the-line Asus, AMD, etc. Unfortunately, AMD lied about the capabilities of the board. They said it was capable of RAID-5 on SATAII, but in actuality it is capable of RAID-5 OR SATAII. Then we installed FreeBSD-CURRENT, and found out it doesn't support RAID-5 yet. Needless to say, it has been a hassle.
So we are debating buying a Sun server and running Solaris on it. My expectation is that there should not be any driver compatibility problems at that point.
So long story not so short -- we are looking at trying to run a bootable RAID-Z (yes, I know that requires some tweaking, but I want to ensure that the boot is also recoverable) with a Xen Dom0. On top of that, we want to allow each domain to run its own OS (whichever they choose, thus prefer AMD-V) as a DomU. Based on this, I think that each DomU would get the advantage of Raid-Z on the underlying filesystem, even if they didn't know about it directly, and regardless of which OS they are running.
Do you have any thoughts, concerns, or questions?
Malachi
A followup:
Hi Malachi,
First of all, sorry to hear about the experiences you've had with the FreeBSD and Asus! Not good.
Secondly, one of the reasons for the Try and Buy promotion is that you can test the capabilities of each machine prior to actually buying it. I'm sure that you will have a positive experience and benefit from this.
I would recommend that we assign a server specialist to you, that can be available prior to your choice of machine, and also be available for you right through the trial process to assist you in your testing, additional components, adjustments in configurations, etc. Would you have a
contact number that we can call you on?Best Regards
Mark Cradock
Sun Microsystems
Tel: +
353
(0)599136768
mark.cradock@your-sun.com
www.sun.com
So the rest will most likely be via phone.
15 January 2007
Chat with agent about Sun Try and Buy Products
You have been connected to Brenda Byrne.
Brenda Byrne: Welcome to Sun Microsystems. How may I help you?
Malachi de AElfweald: Try 2.
Brenda Byrne: Sorry about you getting pushed off a moment ago
Malachi de AElfweald: Hi Brenda. I have a question about Try and Buy products.
Brenda Byrne: Sure - go ahead!
Malachi de AElfweald: I am looking at replacing my server with one to run Solaris (or OpenSolaris), Raid-Z and Xen
Brenda Byrne: ok
Malachi de AElfweald: I assume that the Try and Buy products would all be 100% compatible with Solaris and OpenSolaris, correct?
Brenda Byrne: yes, of course
Malachi de AElfweald: How does the Try and Buy program work?
Brenda Byrne: You can simply receive one of the listed Sun machines on trial, no questions asked, for 60 days, with shipping costs provided by Sun, and at the end of the 60 days, or prior to then, make a choice to keep the machine....or buy it!
Brenda Byrne: Would you be trialing a machine on behalf of a company, or in a personal capacity?
Malachi de AElfweald: Company
Malachi de AElfweald: And do they come with the OS preloaded, or do I install it?
Malachi de AElfweald: Also, is there a subscription fee, or just a one-time cost to purchase?
Brenda Byrne: OS is preloaded
Brenda Byrne: there is a one-time cost to purchase the machines
Brenda Byrne: if I understand you correctly
Malachi de AElfweald: so, for example, it would come with... Solaris10? no subscription/support costs for it?
Brenda Byrne: Yes - they come with Solaris 10 - which is a free operating system anyway.
Brenda Byrne: But no, Solaris 10 customer support does not come provided with the machine
Malachi de AElfweald: available but not required?
Brenda Byrne: There is a fee for signing up to customer support for Solaris 10 (which is not required)
Brenda Byrne: Solaris 10 customer support fees start at $10 per month
Brenda Byrne: What kind of a machine are you thinking of trialing?
Malachi de AElfweald: kk. and is there a way to specify how we want it configured? ie: bootable RAID-Z
Brenda Byrne: yes
Malachi de AElfweald: ok. it all sounds pretty good. I guess at this point I just need to figure out which machine I want.
Malachi de AElfweald: is there anything else I should know?
Malachi de AElfweald: any requirements to qualify or anything?
Brenda Byrne: Nope...except that it is on behalf of your company, and is delivered to your company address
Brenda Byrne: What kind of applications would you test on it?
Malachi de AElfweald: That's fine.
Malachi de AElfweald: A year ago, a polish hacker wiped out our FreeBSD box. everything. multiple domains. I told the FBI, but they didn't even ask what the IP address was.
Malachi de AElfweald: So, I am looking at replacing that server with a new one.
Malachi de AElfweald: Expected usage is that it is going to be a Xen Dom0 and run each domain as a DomU.
Malachi de AElfweald: Primarily, I am a java developer - but that won't really matter since each domain will have its own OS
Malachi de AElfweald: Ok, I think I have all the information I need. Just need to try to decide which machine to go with.
Malachi de AElfweald: Thank you for your assistance.
Brenda Byrne: no problem.
Your session has ended. You may now close this window.

