Showing posts with label gsoc. Show all posts
Showing posts with label gsoc. Show all posts

Thursday, December 15, 2011

Advice for aspiring GSOCers


I have been getting some mails across regarding how to apply to GSOC etc.. I thought it would be better if scribble this piece of advice and mail back the link, it would be easier for me and also for the readers.. :)
So here are a few points on my head, right now.. I will keep editing it so that it retains its relevance. I would happy to get response from you all (good/bad)...


Be clear in what you want out of your summers and what is your field of interest. Once you know yourself, find out the organizations that are working in your field of interest. 


Check out the various projects they have worked in. Find one project of which interests you.
Warning: Don't work for a proposal for many projects. Select one or max two ideas and start tweeking with the code..


Don't get bogged down with a lot of code that you come across, if you are working with an existing project, there are all chances a lot of code is already on place. Even if you are starting out of scratch, clearly jot down the requirements and the approach you are going to follow to meet those. Check with the mentor if she approves of it.


Hang out on irc channel for your org as well as gsoc-in its the indian channel for gsoc participants.. You'd find some good and helpful people there. For project related advice get in touch with previous successful participants for the org, that would be very helpful. and don't worry about the reply of the mentor. They are very busy people and generally come in regular contact only after selection. Don't bug them much, be patient!


A good proposal which captures your thought clearly is the key to all of it. Check out the various proposals submitted by the folks to your organization earlier. You'll get a better idea.
The honeynet project is absolutely an intelligent choice if you are interested in honepots & malwares.
Zhijie Chen was the previous contributor to the project which I worked with. Here are the links he provided me.
http://joyan.appspot.com/tag/phoneyc


Don't be disheartened if its your first time, as Everyone has a first time. So, if you are interested, you MUST apply.
I hope it helps.. 

Wednesday, August 11, 2010

Training the classifier or handling vmware errors!!

Training the classifier doesn't seem to be as much fun as I thought it would be.
Reasons I thought it would be fun:
  1. I had found some new malicious web pages, by simple google searches!!
  2. The lists for training the classifier for both the mal and safe classes was prepared.
  3. The only task was to now to pickle the features dictionary.
Reasons it became a pita:
  1. the vmware error, lack of memory, at the end of 12 URL scan
  2. the continuation of this error now, at each URL, and even after restarting the host machine, and allocating larger RAM to vmware.
  3. finally it rewrote the pickle file that it had learned, means features of 15 URLs..
I have googled for this vmware error, but haven't found any suitable solution.
I thought the OS were trained to handle batch jobs, very early after their birth. This anomalous behaviour is out my understanding!!
Any body with any relevant suggestions??

Sunday, August 1, 2010

Unarchiving Heritrix' archives - arc.gz

Those who have used heritrix for web crawling are aware of  the 'arc' format. For others, Heritrix is a web crawler (a million $ guess :P) which archives the pages it has crawled into arc format.
In order to collect the corpus for training my classifier I thought of using it. Though phoneyc would have been an option, but I had to collect as many samples as possible so I planned on using heritrix.
Configuring it is pretty easy as it has a nice documentation. Well, in order to extract the crawled pages I was searching for some script or tool (I am too foolish and scared of errors while coding, in short i am a noob!), so I wasted quite a lot of time googling.. Sometimes laziness is a boon, I wish I had been lazy to google!
Finally, I mailed Peter Likarish, a Phd. student at University of Iowa, who had previous experience with heritrix and obfuscated JS classification too, and he suggested that arc's are flat files and its pretty easy to extract pages from there. Also, some understanding of sgmllib.py helped me. using a handful of regular expressions and some loops, ta-da!! I got the code up and running!
For interested readers, the code is here.
I have tried it on some arc's, it seems to work fine. Well, in case someone tries to use it and run into a bug, I apologise for their inconvenience. Please let me know in case of problems, bugs, or errors. I would be grateful..

so, start crawling!!!

Tuesday, July 27, 2010

'top' property of 'Window'

The top element of the DOM in Javascript has the following properties:

   1. It refers to the [object window] or the self.
   2. In case of frames or iframes, as obvious, 'top' refers to the top object that created the [i]frame i.e. the [window]
   3. But when we do a window.open(), then the top of this window refer to itself and not to the parent window.
    ps. to refer to the parent window from this newly created window there is the 'opener' property.

The following code explains it clearly...

/************************ Toping.js *******************************/

//var i = "i am at the top."
//function foo(j)
//{
//return eval(3*j);
//}
//alert(foo(3));
//alert("self: "+self.location.href+"\n"+"top: "+top.location.href);
//document.write();
//o = window.open("teesri.js",'win')

/************************ bottom.js ******************************/

//alert(top.i);
//alert("self: "+self.location.href+"\n"+"top: "+top.location.href);
//var foo = top.foo;
//alert("top ka foo in bottom : "+ String(foo(4)));
//window.open("teesri.html","win");

/********************** teesri.js ********************************/

//alert("self: "+self.location.href+"\n top: "+top.location.href);

It might be a trivial concept for some, but it took long to sink in, for me!

Saturday, July 10, 2010

Malicious Javascript - blueprints

Javascript might be a great scripting language but it has been recently been abused a lot to carry out drive-by-downloads attacks. It targets the browsers and the plugins vulnerabilities at the client side.

There are certain features in the structure of the malicious javascript, though, which can be used to detect its presence with high precision. My GSOC 2010 project aims at finding these features and extracting them and thus classify scripts on the basis of these scores into benign and malicious. Finally integrating the complete solution in the low interaction client honeypot - PhoneyC.
I have extracted 9 features, which have been mentioned by a lot of people in their works. These features have been extracted from a very very modest corpus, which is not very broad yet, of 15 benign and 10 malicious JS samples.
The findings expressed as graphs, file against the feature value can be found here. The graphs show malicious scripts features in red and benign scripts in blue.

1. average characters per line
2. average eval() argument length
3. string definition to string use ratio
4. # unicode characters
5. # lines in the script
6. % human readable characters
7. % white space in the script 
8. # words in the script
9. dynamic execution calls

Though the results aren't very encouraging for all the features, but some of them like the string definition to use ratio, % human readable characters, %white space, offer some hope. Improvements in the implementation of the features extraction with little assumptions is required to build a proper extractor for the classifier.
The code for the feature extractor and the classifier may be accessed in my svn branch of phoneyc under njain-anomalydetection.

I sincerely appreciate the comments and reviews on the current work feature extraction and classification.
p.s. - truly speaking,this is my first attempt at regex, pickling, or in short, programming, in that case. All thanks to the mentor for his able guidance and constant motivation.

Monday, June 7, 2010

Anomaly Detection

Anomaly detection is a unique approach to find the odd one out or malicious value. The approach basically involves learning the normal behavior and then detecting variation from this established behavior, which is called a profile. The variation is found based on a model. A model supports in learning as well as detecting. The crux of the approach is "the model".
A basic understanding of the approach can be had from the following program which learns A, an arbitrary integer variable. This model learns that A normally lies between the minimum and maximum values input during the learning mode. It also learns a threshold as 10% of the mean of the entered values. After successfully learning the values of A the model switches to the detection mode. In this mode the difference of the entered value and the mean is compared to the threshold. A difference greater than the threshold is marked anomalous and the value is put in the anomaly list else it is appended to the normal list of values. This is a very naive but working implementation of anomaly detection approach.
The original python implementation is here:

class learna(object):

    def learnA(self):
        """a function to learn a"""
        list=[]

        list=l.learn()
        low=min(list)
        high=max(list)
        avg=sum(list)/len(list)
        print "average is",avg
        l.detect(low,high,avg)


    def learn(self):
        print "learning mode"
        alearned=[]
        for i in range(5):
            al=int(raw_input("enter integer value for a."))
            alearned.append(al)
        lower=min(alearned)
        upper=max(alearned)
        print lower,"<",upper
        return alearned

    def detect(self,low,high,avg):
        print "running in detection mode."
        aentered=[]
        anomaly=[]
        normal=[]
        anomalous=[]
        threshold=0.1*avg
        for i in range(5):
            ae=int(raw_input("enter current integer value."))
            aentered.append(ae)
            if (aehigh):
                anomaly.append(ae)
            else:
                normal.append(ae)
            if (abs(ae-avg)>threshold):
                anomalous.append(ae)
        print "total anomalous value",len(anomaly)
        print "total normal values",len(normal)
        print "total entered values", len(aentered)
        print "total detected anomalous values",len(anomalous)

Sunday, April 11, 2010

Structure of Implementation (original One)

STRUCTURE OF IMPLEMENTATION:
I drew it as an ASCII figure in the proposal and something unusual happened with the whitespaces, thus the figure became a mess. Please use this one for further reference.