Showing posts with label Idea. Show all posts
Showing posts with label Idea. Show all posts

Thursday, August 23, 2018

Draft Tree

Intro

I like blogging about different design patterns, algorithms and approaches to solving different problems. This time I'll describe a problem I ran into some time ago and the approach I found to solve it. I always do not claim to have found a unique and absolutely perfect solution by myself. But I like the idea of keeping and reusing successful and interesting approaches to solving problems. Even if the code cannot be reused the idea itself is often very useful. This is not exactly a design pattern in the classical sense but it is more like a "design approach". But I'll describe it as a design pattern.

Motivation

If one wants to work with a DSL these days he or she usually has to use Antlr. This tool is free (BSD license) but is very useful. The grammars are simple enough. The code can be generated for Java, Javascript and other languages. In fact it is the tool of choice for anyone who wants to take on some DSL-related work.
When the grammar is complete and working well what the user has is the grammar tree. It looks like this:

This example is a logical expression from a simple grammar that parses logical expressions. The tree shows the grammar tree for 1 > 0 OR x < Y AND 10 > Z2. 
Antlr offers great tools to work with these trees: Listeners and Visitors. Listeners fire when a node is encountered and visitors visit nodes one by one. These tools are really great and help a lot. 

But what if the transformation one is trying to accomplish cannot be achieved with Antlr's listeners and visitors? This could be the case if the DSL structure is significantly different from the structure one is trying to achieve (we'll call it 'code structure' from now on. The implication is that the DSL is transformed into source code). 
Some people may try to modify the Antlr tree a little, some may add more complicated listeners/visitors. In this blog post I'll argue that there is a better solution.

Applicability

Use this pattern when you need to perform complicated transformations on the Antlr tree that are hard to accomplish directly with a listener or visitor.

Structure

There is a simple and elegant solution to this problem. We need to introduce an intermediate tree structure. In Java this can be accomplished with the class DefaultMutableTreeNode and other classes from javax.swing.tree. The API may look a bit old-fashioned with Enumeration and no generics but it works. The first tree instance is a copy of the Antlr tree with proper user objects in the nodes. Then we can slowly modify this tree step by step to make it look closer to the code structure. Usually if this structure is used there is no need to do it all in one step. Here is an example of the transformation:

This intermediate tree is like a draft that needs to be improved. After some steps it is ready to be transformed into the final code structure. 

Consequences

  1. Simplicity. The process of transforming the DSL into the code structure can be defined as simple, concise operations on the draft tree. This makes the transformation code less complex. The implementation may even use a modified version of the Assembly Line pattern.
  2. Loose coupling. All the operations on the draft tree are independent of the original DSL and the final code structure. If either the DSL or the code structure has to be modified the draft tree can be adapted to the new form by adding or removing steps. There is not tight coupling between the DSL and the code structure.

Final Thoughts

No sample code can be provided in this case because any code will be more specific than the design approach itself.

Tuesday, August 7, 2018

IDE approach to log analysis pt. 2

Intro

In the first part I explained the theoretical approach to log analysis that I think is best for a sustain engineer. This engineer doesn't need to analyze logs immediately as they come but instead is focused on a deep analysis of complicated issues. In this second part I'll show that many search scenarios can be covered with one sophisticated template and show a working prototype.

Search Object Template

The main requirement for the search template is it must be sophisticated, very sophisticated in the best case. The less manual search the better. A sophisticated template should do most of the work and do it fast. As we don't have any servers here only the developer's PC which is expected to handle 2-3 GB of logs speed is also important. 

Main Regular Expressions

The template should declare some regular expressions which will be searched for (with Matcher.find) in the logs. If more than one is declared first the results for the first are collected, then for the second etc. In the most general sense the result of a search is an array of String - List<String>. 

Acceptance Criteria

Not all results are accepted by the searching process. For example the engineer can search for all connection types excluding "X". Then he or she can create an acceptance criterion and filter them out. by specifying a regex "any type but X". Another possibility is searching within a time interval. The engineer can search for any log record between 10 and 12 hours (he or she has to enter the complete dates of course). 
Looking for distinct expressions is also possible. In this case the engineer specifies one more regular expression (more than one in the general case). An example will explain this concept better. 
distinct regex: connection type (q|w)

log records found by the main regex:
connection type w found
connection type q created
connection type s destroyed
connection type q found

The result of a distinct search:
connection type w found
connection type q created

Parameters

One of the issues with regular expressions is that really useful regular expressions are very long and unwieldy. Here is a sample date from a log:
2018-08-06 10:32:12.234
And here is the regex for it:
\d\d\d\d-\d\d-\d\d \d\d:\d\d:\d\d.\d\d\d
The solution is quite simple - use substitution. I call them parameters for the regex. Some parameters may be static like the time for the record but some may be defined by the user. Immediately before the execution the parameters are replaced with the actual values.

Views

The result of the search is a log record i.e. something like
2018-08-06 10:32:12.234 [Thread-1] DEBUG - Connection 1234 moved from state Q to state W \r?\n
While it is great to find what was defined in the template it would be even better to divide the information into useful pieces. For example this table represents all the useful information from this record in a clear and concise way:

Connection 1234 Q -> W

To extract this information pieces we can use the "view" approach. This means declaring smaller regexes that are searched for in the log record and return a piece of information about the log record. It is like a view of this log record. Showing it all in a table makes it easier to read. Also a table can be sorted by any column.

Sort and Merge

The most efficient way to make this kind of search with the template is use a thread pool and assign every thread to a log file. Assuming there are 3-4 threads in the pool the search will work 3-4 times faster. But merging results becomes an important issue. There can be 2 solutions here:
  1. Merging results. We need to make sure that the results go in the correct order. If we have 3 log files, the first one covering 10-12 hours, the second 12-14, the third 14-17 then the search results from those file must go in the same order. This is called merging.
  2. Sorting results. Instead of merging them we can just sort them by date and time. Less sophisticated but simple.
Merging looks like a more advanced technique which allows us to keep the original order of records.

Workflow


Final Thoughts

The question that must be nagging everyone who has reached this point in this post is: Has anyone tried to implement all this? The answer is yes! There is a working application that is based on the Eclipse framework, includes a Spring XML config and a lot of other stuff. The search object templates work as described in this article.
Here is the Github link:
https://github.com/xaltotungreat/regex-analyzer-0
Why 0? Well it was meant to be a prototype and to some extent is still is. I called this application REAL
Regular
Expressions
Analyzer
for Logs
It is assumed the user has some knowledge how to export an Eclipse RCP application or launch it from within the Eclipse IDE. Unfortunately I didn't have enough time to write any good documentation about it. By default it can analyze HBase logs and there are a lot of examples in the config folder. 

Tuesday, July 31, 2018

IDE approach to log analysis pt. 1

Intro

I think most software engineers understand the importance of logs. They have become part of software development. If something doesn't work we try to find the cause in the logs. This could be enough for simple cases when a bug prevents an application from opening a window. You find the issue in the logs, look it up on Google and apply the solution. But if you are fixing bugs in a large product with many components analyzing logs becomes the main problem. Usually sustain engineers (who are fixing bugs not developing new features) need to work with many hundreds of megabytes of logs. The logs are usually split into separate files 50-100 MB each and zipped. 

There are several approaches to make this work easier. I'll describe some existing solutions and then explain a theoretical approach to this problem. This blog post will not discuss any concrete implementations. 

Existing Solutions

Text Editor

This solution is not actually a solution it is what most people would do when they need to read a text file. Some text editors could have useful features like color selection, bookmarks which can make the work easier. But still the text editor falls short of a decent solution.

Logsaw

This tool can use the log4j pattern to extract the fields from your logs. Sounds good but these fields are already obvious from the text. Clearly the improvement is insignificant over a simple text editor. 

LogStash

This project looks pretty alive. But this approach is quite specific. Even though I've never worked with this tool from the description I understood that they use ElasticSearch and simple text search to analyze logs. The logs must be uploaded somewhere and indexed. After that the tool may show the most common words, the user may use text search etc. Sounds good, seems to be some improvement. Unfortunately not so much. Here are the cons:
  • Some time is required to begin working with the logs. One has to upload them, index them. After the work is done these logs must be removed from the system. Looks like a little overkill if the logs are meant to be analyzed and discarded. 
  • A lot of components involved with a lot of configuration necessary.
  • Full text search is not very useful with logs. Usually the engineer is looking for something like "connection 2345 created with parameter 678678678". Looking for "created with parameter" will return all connections. Looking for "connection 2345" will return all such statements but usually there is only one - when this connection was created.

Other Cloud-based Solutions

There are a lot of cloud-based solutions available. Most of them have commercial plans and some have free plans. They offer notifications, visualizations and other features but the main principles are the same as for LogStash. 

Log Analysis Explained 

To understand why these solutions do not perform well for analyzing complex issues we need to try to understand the workflow. Here is a sample workflow with the text editor:
  1. An engineer received 1 GB of logs with the information that the bug happened at 23:00 with request ID 12345.
  2. First he or she tries to find any errors or exceptions around that time.
  3. If that fails the engineer has to reconstruct the flow of events for this request. He or she begins looking for statements like "connection created", "connection deleted", "request moved to this stage" trying to narrow down the time frame for the issue. 
  4. That is usually successful (even though could take a lot of time) now it is clear that the issue happened after connection 111 was moved to state Q. 
  5. After digging a little more the engineer finds out that this coincides with connection 222 moving to state W. 
  6. Finally the engineer is delighted to see that the thread that moved connection 222 to the new state also modified another variable that affected connection 111. Finally the root cause.
In this workflow we see that the engineer most of the time is looking for standard strings with some parameters. If only it could be simplified...

IDE approach

There are several parts to the IDE approach.
  1. Regular expressions. With regular expressions one can specify the template and search for it in the logs. Looking for standard strings is much simpler with regular expressions. 
  2. Regular Expressions Configuration. The idea here is that standard strings like "connection created \d{5}\w{2}", "connection \d{5}\w{2} moved to stage \w{7}", "connection\d{5}\w{2} deleted" do not change often. Writing the regular expression to find it every time is unwieldy because such regexes could be really long and complicated. It is easier if they can be configured and used by clicking on a button.
  3. IDE. We need some kind of an IDE to unite this together. To read the configuration, show the logs files and stored regexes, display the text and search results. Preferably like this:
  4. Color features. From experience I know that log analysis is much easier when you can mark some strings with color to easily see it in the logs. Most commercial log analyzer tools use color selection. The IDE should help with that.

Pros and Cons

Pros of the IDE approach:

  1. No cloud service necessary. No loading gigabytes of logs somewhere, no cloud configuration. One only has to open the IDE for logs, open the log folder and start analyzing.
  2. If the IDE is free the whole process is completely free. Anyway should be cheaper than a log service.

Cons of the IDE approach:

  1. Most cloud services offer real-time notifications and log analysis "on the fly". It means as soon as the specified exception happens the user is notified. The IDE approach cannot do that.
  2. The requirements for the user's PC are somewhat higher because working with big strings in Java consumes a lot of memory. 8 GB is the minimum requirement from my experience. 
The bottom line is the IDE approach is suitable to analyze complicated issues in the logs. It cannot offer real-time features of cloud services but should be much cheaper and easier for analyzing and fixing bugs. 

Final Thoughts

It would be great if someone could implement this great approach! I mean create this IDE with all those features and make log analysis easier for everyone! I know from experience that this could be a tedious work that feels harder than it actually is. In the next post (part 2) I'll explain the difficulties/challenges with this approach and offer a working implementation based on the Eclipse framework.

Thursday, July 26, 2018

On the question of licenses

Intro

Hi all. Today's topic is indirectly related to Java. I think it is not an overestimation that most people hate software licenses. I'm one of them. Every time I begin reading a license it makes my brain go crazy. Like this excerpt from the Gnu Gpl:
You may make, run and propagate covered works that you do not convey, without conditions so long as your license otherwise remains in force. You may convey covered works to others for the sole purpose of having them make modifications exclusively for you, or provide you with facilities for running those works, provided that you comply with the terms of this License in conveying all material for which you do not control copyright. Those thus making or running the covered works for you must do so exclusively on your behalf, under your direction and control, on terms that prohibit them from making any copies of your copyrighted material outside their relationship with you.
 All these "covered but not limited but complying with..." and "on your behalf or on your protection or ...". It must take several years to get comfortable with this language. 95% of the people in the world are not legal experts and cannot understand most of it. This is the main reason why these unpleasant things may happen. People may agree to anything because most of them do not read the terms and they do not read them because they are hard to understand.

I read some blog posts suggesting that AI may help with this. Like it will be scanning licenses and contracts for "bad" parts. The person who created this idea must have really disliked AI and AI developers. Even courts and judges have difficulties trying to understand the contracts and licenses and he or she wanted AI to handle it.

Solution

What is funny this problem does have a fairly easy solution. And this solution is well-known. Most programmers and also java programmers use this solution every day. It is (tense music in the background) INHERITANCE. Yes simple plain old inheritance. Even Java-style inheritance will do. Licenses must get rid of the old model when every license contains all the provisions, information that is necessary to use it. Licenses for concrete products must inherit well-known licenses like Java classes. So instead of the brain-depressing text above the Gnu Gpl could look like this:
If someone is presented with this structure he or she can clearly see everything he or she needs. And the mind-boggling text is not necessary. If someone wants to use the product he or she sees it is free (inherited from "Free Software License" which means no payment necessary). If someone wants to use it as part of a commercial product he or she sees that this license is non-commercial (inherited from "Non-commercial free software license") and cannot be used for this purpose. If someone wants to know the specifics he or she can read any of these texts to find out.

Explanation

With this scheme everybody gets what they want. The legal experts still have those mind-boggling texts. The texts do not go away and the current legal system doesn't need to change their ways much by incorporating AI to read the texts with them. The users get a clear explanation of the license they can understand.

This scheme is based on some assumptions. First there should be well-defined "abstract" licenses like abstract Java classes which define the basic terms. For example there could be an abstract license for  proprietary software and free software. These abstract licenses are organized in hierarchies. The provisions of the licenses higher in the hierarchy cannot be overridden by the licenses lower in the hierarchy. It is like all the methods/fields of the abstract classes are final. If two provisions contradict each other the provision from the license higher in the hierarchy is used.

Sounds really simple. Why has nobody used this solution before? I don't know. For unknown reasons the judicial system uses the one-license-has-all approach. It is like they work in Java without inheritance always writing the same code even though they need to change only one line. The people in the software world at least understand why writing the same code many times is bad and they came up with the idea of standard licenses like GNU GPL, Apache which people can use if they need it. While it is a good idea it falls short of a comprehensive solution to this problem.
Inheritance could be that solution.