Sunday, 2 December 2007

And Finally

This may be my last blog entry as I am just completing the last question of TMA04. The poster is done. I hope it is OK but it's too late to worry about now. I did it on pre-processing which may not be technical enough.

Anyway that's it. Top Gear is on in 30 mins and I'm off.

Sunday, 11 November 2007

Almost There

Looks nice. Have no idea how good it is. I'm not that impressed. I keep thinking about things I could have done. Will print out final version this evening

Saturday, 3 November 2007

Saturday 3rd November

Continuing to write up. Taking longer than I would like

Monday, 22 October 2007

Monday 22nd October

Last night I finished all the practical work. I managed to test a MLP with the same inputs as the KBS. I tried various certainty factors and rules with the KBS to improve performance but it could not match the MLP. At least this is a direct comparison of KBS vs NNs which all the other experiments are not.

Began write up today. I will just carry on from the draft really as there wasn't much wrong with that. Word 2007 does make it look impressive. So even if it's crap it will look good.

Sunday, 21 October 2007

Sunday 20th

Rechecked scoring of KBS experiments and corrected a few mistakes. Both experiment scoring the same though PPVs are different.

A few observations about these KBS:

Is there any point in putting three rules in for the grade as by inference if it’s not one then it’s one of the other two? Would two rules be enough

There are no cases in the training set with 18+ lymph nodes but this rule does exert an influence on the other cases. Do I remove it?

Currently modifying original KBS to test this.

----------------
Now playing: The Avalanches - Frontier Psychiatrist
via FoxyTunes

Saturday, 20 October 2007

Saturday 19th October

A rare weekend day with not a great deal to do. So more OU. Not much more to do now. I completed testing and scoring my flex program yesterday. I built a MLP last night and tested the same training file as I have been using to test the flex program, however I did this with the full 32 input neurons. The MLP scored very well but it will be interesting to see how it scores with the same number of inputs as the flex program. This is what I am doing now. I am modifying the KBS input files for the MLP.

I have in the mean time modified the certainty factors in the flex program to try to improve the performance but without success.

Will complete this practical work and begin writing up the results and the project this weekend.

T-25 days

----------------
Now playing: Super Furry Animals - Lazer Beam
via FoxyTunes

Friday, 19 October 2007

Friday 19th Arrrggghhhh!

Tried copying the line "prognosis is reccurence" from rule r1c to everywhere it occured in the other rules etc. Then recompiled and tried again. Rule r8c now working properly. Cannot see why this happened. See below:

C.F. : TRY : r8c
C.F. : LOOKUP : (grade is '3') = -1
C.F. : IMPLIES : cf(0.35) @ -1 -> -0.35
C.F. : LOOKUP : (prognosis is reccurence) = 0.2
C.F. : CONFIRMS : 0.2 + -0.35 -> -0.1875
C.F. : UPDATE : (prognosis is reccurence) = -0.1875
C.F. : FIRED : r8c

Looked at the certainty factor example in Chapter 3 of T396 and it doesn't simply add certainty factors together.

Start testing KBS again!

----------------
Now playing: Maximo Park - Going Missing
via FoxyTunes

Friday

Yesterday I did a number of experiments with NNs. I couldn't get a score of more than 51% for any network topography despite adjusting training cycles etc. The performance of many of the networks was remarkably similar.

I identified those cases which gave false +ve and -ve results and added them to the training sets but performance deteriorated. The data is ambiguous. Some patients have virtually the same sets of data but different outcomes. There are not enough data items to improve the performance - no ER/PR results etc.

Preprocessed the training set for the KBS. I couldn't get excel to output a .csv file so I had to produce it manually. I then copied these into word so that I could paste who strings into the console window of flex.

Today I have been running these and have noticed that rule 8c is not working properly as it is replacing the certainty from the previous rules and replaces it.

C.F. : TRY : r9c
C.F. : LOOKUP : (location is central) = -1
C.F. : LOOKUP : (grade is '1') = 1
C.F. : AND : -1 + 1 -> -1
C.F. : IMPLIES : cf(-0.5) @ -1 -> 0.5
C.F. : LOOKUP : (prognosis is reccurence) = 0.1
C.F. : CONFIRMS : 0.1 + 0.5 -> 0.55
C.F. : UPDATE : (prognosis is reccurence) = 0.55
C.F. : FIRED : r9c

C.F. : TRY : r8c
C.F. : LOOKUP : (grade is '3') = -1
C.F. : IMPLIES : cf(0.35) @ -1 -> -0.35
C.F. : UPDATE : (prognosis is recurrence) = -0.35
C.F. : FIRED : r8c

C.F. : TRY : r7c
C.F. : LOOKUP : (grade is '2') = -1
C.F. : IMPLIES : cf(0.1) @ -1 -> -0.1
C.F. : LOOKUP : (prognosis is recurrence) = -0.35
C.F. : CONFIRMS : -0.35 + -0.1 -> -0.415
C.F. : UPDATE : (prognosis is recurrence) = -0.415
C.F. : FIRED : r7c

I need to sort this before continuing.

Also if the CF.:CONFIRMS lines are supposed to be adding the numbers together they are wrong.

C.F. : TRY : r11c
C.F. : LOOKUP : (size is less_than_15) = 1
C.F. : IMPLIES : cf(-0.25) @ 1 -> -0.25
C.F. : LOOKUP : (prognosis is reccurence) = -0.5
C.F. : CONFIRMS : -0.5 + -0.25 -> -0.625 [-0.5+ -0.25 = -0.75]
C.F. : UPDATE : (prognosis is reccurence) = -0.625
C.F. : FIRED : r11c

Wednesday, 17 October 2007

Wednesday 17th October 11p.m.

A day of testing NNs. All the MLPs I tested regardless of the number of training iterations or hidden layer neurons gave very similar results. So I analysed the results of Exp E, as it was fairly representative, for false +ves and -ves. I then produced a new training file of some of these and the original random cases but still with 50 cases.

The results of an MLP with 10 hidden layer neurons and 50000 training iterations were poorer than using the random set.

Next try a training set of all false +ves and -ves plus orginal training set.

Wednesday 17th October 11a.m.

More experimenting. Three days to complete the experiments and then a couple of weeks to write up. Well that's the plan anyway.

First up complete neural network experiments. Not happy with the ones I've carried out before they are to ad hoc. I will organise them better this time.

So first create an MLP and work out the optimum number of training cycles. Do each experiment at least twice because NNs do not always cluster data the same way each time they are trained.

Score with score tool and calculate sensitvity, specificity and PPV.

Then experiment using different numbers of hidden layer neurons to find optimum.
Then, well I'll see how it's going.

----------------
Now playing: Jeff Wayne - Dead London
via FoxyTunes

Saturday, 22 September 2007

22nd September.

Tried a Kohonen SOM during the week with the defaults from NeuralWorks and using the entire 285 cases. With the score tool it score 51% which is higher than any MLP but those were only trained with the training set. I will rerun both networks later to ensure I'm comparing like with like. Then I'll compare this with my KBS.

Tuesday, 4 September 2007

4th September

Well time to get on after a couple of weeks doing sod all. Really surprised and pleased by the result for TMA03. I'm not entirely sure which task to do next. I'll probably do some more preprocessing of the data to enable me to test the flex program and score the output. I think this will be quite labourious so I'd better do it sooner rather than later.

Wednesday, 15 August 2007

Wednesday 15th August

I'm making this decision as I type it into TMA03. I mentioned on the blog before that I wanted to compare like with like so I wanted to construct a neural network with the same number of inputs as the flex program. So for a rule inthe flex program such as:

uncertainty_rule r4c

if the involved_nodes is '>=6 and <=17'

then the prognosis is reccurence

with certainty factor 0.20 .

The equivalent input in the neural network will be given by whether the statement:

Involved node is greater than or equal to 6 but less than or equal to 17

is true or false.

Monday, 13 August 2007

Monday 13th August - Oops

Noted when writting up the draft for TMA03 that I have not mentioned that I decided to use Positive Predictive Value, Specifity and Sensitivity as performance indicators as these are more universally understood than the OU score tool.

Monday 13/08/07

Spent most of yesterday writing Q2 of TMA03. I'm not sure I entirely understand what is menat by the "doing" part of the project but I've done what I can, bearing in mind the practical work is far from repeat.

I've actually started to enjoy using Flex. I've been modifying the program that uses certainty factors and added more rules (kbs 6 and 7) and have tested with a couple of patients and it seems to work fine. The certainties of the evidence which are used at the start of the program have caused much head scratching. For example:

All patients under 30 years old do not suffer recurrence. So I can asign a certainty of -1.0 to the rule for the certainty that the patient will suffer recurrence. But if the patient is over 30 the evidence then the statement "is the patient under 30" is definitely not true so the certainty factor in the starting statement entered in the console would be -1. Implying that recurrence must occur which isn't true, I think.

So I've fiddled with the certainty factors in the rules to try to overcome this problem and tested with -1 or 0 in the starting statement when the condition is not true.

I am actually finding the TMA a pain because I actually have some enthusiasm for Flex at the moment and want to get on, but I have to finish the TMA.

Saturday, 11 August 2007

Saturday 11th August

Oh well so much for the new season. Some things don't change. Found a useful, if rather old reference for prognostic indicators in breast cancer.

Histopathology

Volume 19 Issue 5 Page 403-410, November 1991

To cite this article: C.W. ELSTON, I.O. ELLIS (1991)
pathological prognostic factors in breast cancer. I. The value of histological grade in breast cancer: experience from a large study with long-term follow-up
Histopathology 19 (5), 403–410.

Need to look for it at work.

Friday, 10 August 2007

Friday 11th August

Modified flex certainty factor program. Program compiles and runs but the results are not as expected. The certainty factor associated with the prognosis is adjusted after each rule so that the certainty factor is the sum after all the rules are evaluated but if all patients under 30 do not have recurrence it doesn't matter what the values the rules which look at location, no of LNs etc assign the certainty that will always be 1 however the other rules are affecting this value. Not easy to explain.

Will read through Hopgood and the example in block one of T396. I think I will need less rules with or staements in them.

TMA really needs to get moving tomorrow. Not sure how to tackle it. Too late to contact tutor now. Will do the best I can in the time I have.

Must state that all I want is to pass this project. After 7 years of OU I am utterly fed up with studying.

Football season kicks of tomorrow. The only time I will be away from this PC will be to watch the Hammers.

Wednesday, 8 August 2007

Wednesday 8th August

TMA deadline looming. Need to get it sorted.

Spent some time modifying my flex program using certainty factors. It is almost working. Hopefully a little more time will sort it out. I will see if I can add to the 6 rules I have. If not I will test when it is working. Need to sort out testing and scoring strategy. Probably use a similar score tool as the NN but with less instances to test. Upto 50 patients would be enough.

Then I'll modify the NN so that it has the same data inputs as the flex program and test. That way I'll be testing like with like.

Probably resume this work now after the TMA.

Sunday, 5 August 2007

Later Sunday

Too hot to be really productive. However question 1 of TMA03 is OK. Not entirely sure about question 2.

Decided just to run the flex program with uncertainty rules only as I am sure I can get it to work, though it is cumbersome to test.

Noticed in flex manual that it states:

Given a rule:
rule1: if A & B then C
there are 3 potential areas for uncertainty.
- Uncertainty in data (how true are A and B)
- Uncertainty in the rule (how often does A and B imply C)
- Impreciseness in general
The first 2 can be handled using probabilities and the third using fuzzy logic.

As the main sources of uncertainty in this project is in the data - difficult to measure tumours accurately etc, and the rules, there are only four rules that are always true according to the statistical analysis of this dataset but these are probably not always true with other datasets. Therefore I was justified in using uncertainty rules to deal with these sources of uncertainty.

Sunday 5th August

Back from holiday. I did think I might be able to work on the laptop while away but never did. So am going to work some more on the project draft for TMA03. I have decided that I must compare like with like. In the project I did for T396 there were many more inputs and bird instances used in the neural network than the KBS. Therefore it was hardly surprising it out performed the KBS. After I have a satisfactory NN and KBS working I will either scale up the KBS to use the same number of patient instances as the NN and data inputs, or more likely I will scale down the NN to use the same number of patient instances as the KBS etc.

Back to work.

About Me

My goal in life is to become grumpier. There's no point getting older unless you become grumpier. Working for the NHS helps as does supporting West Ham, so one day I'll end up like Victor Meldrew.