Showing posts with label python. Show all posts
Showing posts with label python. Show all posts

Tuesday, March 15, 2011

Python Boto + SSH SOCKS = Easy Cheap Coffeeshop VPN

A while back (well actually after reading Unencrypted Wifi Must Die) I started using SSH SOCKS tunnels to a small AMI I created with using the AWS Console. I end up paying less than a dollar a month for this. You'll need to install Python Boto for this, but on Ubuntu 10.04 and later this is as simple as

apt-get install python-boto


The code is pretty simple

First connect to EC2

e = boto.connect_ec2()


Create a new instance where AMI_TYPE is set to 't1.micro' (the smallest cheapest AMI)

r = e.run_instances(AMI,key_name=KEY_NAME,instance_type=AMI_TYPE)
i = r.instances[-1]


You basically have to create the reservation and then pull the instance from the list of reservations. Wait for the instance to come up so you can find out the dns name

print "Host name:",i.public_dns_name

--
ec2unnel.py is available on GitHub. You obviously need to change the name to your SSH key and set your AWS environment variables or hard code them in the script.


mfranz@mfranz-xp60sublts:~$ ./ec2unnel.py stop
Connecting to EC2...
Host name: ec2-75-101-195-59.compute-1.amazonaws.com
State: shutting-down
State: shutting-down
State: shutting-down
State: shutting-down
State: shutting-down
State: shutting-down
State: terminated
mfranz@mfranz-xp60sublts:~$ ./ec2unnel.py start
Connecting to EC2...
Creating instance in: us-east-1
Launch time: 2011-03-15T12:15:05.000Z
State: pending
State: pending
State: pending
State: pending
State: pending
State: running
Host name: ec2-50-17-19-144.compute-1.amazonaws.com

mfranz@mfranz-xp60sublts:~$ ssh -v -i duo.pem -D 1080 ec2-user@ec2-50-17-19-144.compute-1.amazonaws.com
OpenSSH_5.3p1 Debian-3ubuntu5, OpenSSL 0.9.8k 25 Mar 2009
debug1: Reading configuration data /etc/ssh/ssh_config
debug1: Applying options for *
debug1: Connecting to ec2-50-17-19-144.compute-1.amazonaws.com [50.17.19.144] port 22.
debug1: Connection established.
debug1: identity file duo.pem type -1
debug1: Remote protocol version 2.0, remote software version OpenSSH_5.3
debug1: match: OpenSSH_5.3 pat OpenSSH*
debug1: Enabling compatibility mode for protocol 2.0
debug1: Local version string SSH-2.0-OpenSSH_5.3p1 Debian-3ubuntu5
debug1: SSH2_MSG_KEXINIT sent
debug1: SSH2_MSG_KEXINIT received
debug1: kex: server->client aes128-ctr hmac-md5 none
debug1: kex: client->server aes128-ctr hmac-md5 none
debug1: SSH2_MSG_KEX_DH_GEX_REQUEST(1024<1024<8192) lang =" en_US.utf8">


And you see it takes a while for the terminated instances to go away.


mfranz@mfranz-xp60sublts:~$ ./ec2unnel.py
Connecting to EC2...
2011-03-15T12:02:23.000Z terminated
2011-03-15T12:15:05.000Z terminated


Obviously you'll need to set your proxy

Sunday, February 13, 2011

Large Scale Packet Dump Analysis with MongoDB

So when dealing with hundreds of MBs (or even several GBs) of packet captures spread across dozens of files, wireshark sort of breaks down, even with a fast CPU. In my case I have a laptop (with LUKS encrypted drive) so it is pretty slow. Yeah you can split them into smaller files but then you lose visibility into the complete picture when performing your queries. I think you could also write some .lua wireshark but you still have the bottleneck of tshark. So what to do? Let's back up a bit.

Over the years I've written a variety of Perl, Python, or Ruby scripts for processing .pcaps. Some that use C (or pure Python) .pcap parsers--or when I first started doing this over a decade ago just parsing the output of tcpdump and building hashes or dictionaries. Not only is this slow but you have to persist your hashes via pickling. And the challenge of the pcap libraries is they typically don't have any application layer decoding and they require a C version of the library which isn't very cross platform.

Enter pdml. I first described this in a Digital Bond blog post back in 2006. I can remember doing this on a lowly Powerbook G4 and it just worked. A year ago I was looking for a project to use MongoDB. So I wrote some code to automate the process of creating the .pdml files using wireshark and extracting the fields of interest and inserting them into a MongoDB database. I have a configuration file that specifies which PDML fields I want to extract.


[frame.len]
type = decimal

[ip.ttl]
size = 8
type = decimal

[ip.src]
size = 4
type = ipaddr

[ip.dst]
size = 4
type = ipaddr

[tcp.dstport]
size = 16
type = decimal



Because MongoDB doesn't support periods in key names (I learned this the hard way last year) I change the name of the field from ip.src to ip_src. Anything that wireshark knows about I can extract and it will become a key for that packet.

While I was still importing packets (I had close to 2 million packets in the database) I issue a query to see the unique source IPs. I can do this for any field that is supported in the PDML.


>> db.raw.distinct("ip_src")

Sun Feb 13 11:20:05 [conn17] query art0.$cmd ntoreturn:1 command: { distinct: "raw", key: "ip_src", query: {} } reslen:4457 1607ms


Or let's say I wanted to look at what are the unique TTLs.


>> db.raw.distinct("ip_ttl")

11:31:17 [conn17] query art0.$cmd ntoreturn:1 command: { distinct: "raw", key: "ip_ttl", query: {} } reslen:452 1779ms


Not bad for 2.7 million packets (and counting).

This example is pretty uninteresting because it is just standard TCP/IP headers and if you just wanted session data you could just use netflow but this far more flexible.

The downside is speed of import. Creation of the .pdml file by running tshark is very slow and the parsing of XML in Python is also not the speediest. I'm up to close to 3.1 million packets in about an hour that I've successfully imported into my database, but once they are in it is lightning fast and you are free. Where I'm (hopefully) headed today is some scripts that will create Graphviz representations of all the communications of interest perhaps like those available with Afterglow. Or I can use this to analyze and reconstruct streams from some of the proprietary protocols that I was most interested in. Or I can use this as an exercise to write a Node.js app to browse this data. The point is getting it into a useful database that allows flexible and fast queries and offloads a lot of manual tasks I would normally have to do.