Skip to content
New issue

Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.

By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.

Already on GitHub? Sign in to your account

Added Microsoft.ML.Benchmarks Project #62

Merged
merged 10 commits into from
May 15, 2018
Merged

Added Microsoft.ML.Benchmarks Project #62

merged 10 commits into from
May 15, 2018

Conversation

KrzysztofCwalina
Copy link
Member

@KrzysztofCwalina KrzysztofCwalina commented May 7, 2018

Added first demo benchmark. We should update it later to something more representative.
This is the first step in resolving #20

To run the benchmark:

  1. Build Release
  2. Open command window in tests\Microsoft.ML.Benchmarks folder
    • if you want to measure hardware counters (not yet enabled), open admin console.
  3. dotnet run -c Release Microsoft.ML.Benchmarks.csproj
  4. Select benchmark to run from the menu

Note: currently, LearningPipeline logs to the windows console by default, which impacts the performance run. You can minimize the overhead by minimizing the console window while it runs the test.

Here is the output from the run:

BenchmarkDotNet=v0.10.14, OS=Windows 10.0.16299.431 (1709/FallCreatorsUpdate/Redstone3)
Intel Core i7-6700 CPU 3.40GHz (Skylake), 1 CPU, 8 logical and 4 physical cores
Frequency=3328124 Hz, Resolution=300.4696 ns, Timer=TSC
.NET Core SDK=2.1.300-rc1-008662
  [Host] : .NET Core 2.0.7 (CoreCLR 4.6.26328.01, CoreFX 4.6.26403.03), 64bit RyuJIT

Toolchain=InProcessToolchain  LaunchCount=1  TargetCount=3  
WarmupCount=3  
Method Mean Error StdDev AccuracyMacro Gen 0 Gen 1 Gen 2 Allocated
TrainIris 215,636,477.31 ns 39,847,380.894 ns 2,251,450.9146 ns 0.98 8000.0000 250.0000 62.5000 11797870 B
PredictIris 1,027,876.15 ns 393,249.160 ns 22,219.3068 ns 0.98 27.3438 13.6719 1.9531 90408 B
PredictIrisBatchOf1 12.53 ns 2.905 ns 0.1641 ns 0.98 0.0171 - - 72 B
PredictIrisBatchOf2 12.83 ns 4.585 ns 0.2590 ns 0.98 0.0171 - - 72 B
PredictIrisBatchOf5 14.04 ns 2.643 ns 0.1493 ns 0.98 0.0171 - - 72 B

@shauheen shauheen requested a review from markusweimer May 7, 2018 23:59
Microsoft.ML.sln Outdated
@@ -85,100 +85,198 @@ Project("{2150E333-8FDC-42A3-9474-1A3956D46DE8}") = "specs", "specs", "{2DEFC784
Documentation\specs\mvp.md = Documentation\specs\mvp.md
EndProjectSection
EndProject
Project("{9A19103F-16F7-4668-BE54-9A1E7A4F7556}") = "Microsoft.ML.Benchmarks", "test\Microsoft.ML.Benchmarks\Microsoft.ML.Benchmarks.csproj", "{77705689-F08D-44B5-A775-3F844EE744AC}"
Copy link
Member

@markusweimer markusweimer May 8, 2018

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shall we mark the .sln as binary? Merging it is usually hopeless anyway :) #Closed

Copy link
Member Author

@KrzysztofCwalina KrzysztofCwalina May 8, 2018

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@eerhardt/@danmosemsft, any opinions on this? #Closed

Copy link
Member

@eerhardt eerhardt May 8, 2018

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is possible to merge .sln files. They are rather touchy though - white space is actually important.

We don't do this in any other repo that I'm aware of. And it is always good to be able to read/review the .sln file edits. So unless we have a compelling reason to make this change, I'd vote we keep it as-is. #Closed

Copy link
Member

@danmoseley danmoseley May 9, 2018

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Text works ok in other repos. #Closed

Copy link
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SLN is a text file. A messy one perhaps. It has been merged as a text file for at least 10 years. On few occasions it was quite useful to see the text. With proper training it is quite mergeable


In reply to: 186587605 [](ancestors = 186587605)

Microsoft.ML.sln Outdated
{46F2F967-C23F-4076-858D-33F7DA9BD2DA}.Debug|Any CPU.ActiveCfg = Debug|Any CPU
{46F2F967-C23F-4076-858D-33F7DA9BD2DA}.Debug|Any CPU.Build.0 = Debug|Any CPU
{46F2F967-C23F-4076-858D-33F7DA9BD2DA}.Debug|x64.ActiveCfg = Debug|Any CPU
Copy link
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should not have any configurations for Any CPU. ML.NET only supports x64 as of now.

Copy link
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that AnyCpu is a preexisting (before this PR) configuration. I was actually trying to remove it, but VS gets confused and adds it back. I can play a bit more with it, but I think it should be a separate PR as it's a separate/preexisting issue).

<PropertyGroup>
<OutputType>Exe</OutputType>
<LangVersion>7.2</LangVersion>
<TargetFramework>net472</TargetFramework>
Copy link
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reason this doesn't target .NET Core or .NET Standard?

Copy link
Member Author

@KrzysztofCwalina KrzysztofCwalina May 8, 2018

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's an exe, so it cannot be .NET Standard (which is for dlls). It could be .NET Core app, but NetFx app makes it easier to run it, as the project output is an exe.
Also, the output of this project does not really matter: BenchmarkDotNet generates the actual program that is being run during the test, and the generated project can be .NET Core (it's just a change in the config I can make if we prefer .net core test).

Copy link
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Still, the dependency of .NET 4.7 makes this code unrunable on Linux, right? Can't we make it dual-target? netcoreapp and net472?

Copy link
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Makes sense. I will change the host app to .NET Core too.

{
[Benchmark]
[MemoryDiagnoser]
public void Iris()
Copy link
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it possible for us to mark a bunch of integration tests as benchmarks in this way?

Copy link
Member Author

@KrzysztofCwalina KrzysztofCwalina May 8, 2018

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

They would need to be a part of an exe project (like this one). But maybe we can try to expose the integration tests in public APIs (of the test project) and then just run them from a benchmark stub in this project. @adamsitnik? #Pending

Copy link
Contributor

@glebuk glebuk May 10, 2018

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Now there are at least two possible ways to do this: Create those items as unit tests (comment above) OR, use a general-purpose command line tool to execute specific pipelines.
For the second scenario - command line tool:
It seems that such command tool would be great to have in contexts outside of benchmarking. Does it really make sense to create a benchmarking-only tool?
Why not simply port a MaML.EXE and create a bunch of *.RSP files to define each benchnarks? That would save us a lot of time developing this AND would give users a CLI for ML.NET
Opened issue #108 to track this.


In reply to: 186589537 [](ancestors = 186589537)

@mmitche
Copy link
Member

mmitche commented May 8, 2018

@dotnet-bot test this please

1 similar comment
@mmitche
Copy link
Member

mmitche commented May 8, 2018

@dotnet-bot test this please

@mmitche mmitche closed this May 8, 2018
@mmitche mmitche reopened this May 8, 2018
@markusweimer
Copy link
Member

Do we have an issue for this? I think we should capture and discuss the intentions of this added project somewhere.

@KrzysztofCwalina KrzysztofCwalina changed the title Added Microsoft.ML.Benchmarks Project [NO MERGE] Added Microsoft.ML.Benchmarks Project May 8, 2018
@KrzysztofCwalina KrzysztofCwalina changed the title [NO MERGE] Added Microsoft.ML.Benchmarks Project Added Microsoft.ML.Benchmarks Project May 8, 2018
@KrzysztofCwalina
Copy link
Member Author

KrzysztofCwalina commented May 8, 2018

Added AccuracyMacro stat to the report. Not sure if this is the best way @adamsitnik? and not sure if this is the best stat @glebuk?

@KrzysztofCwalina
Copy link
Member Author

KrzysztofCwalina commented May 8, 2018

Do we have an issue for this? I think we should capture and discuss the intentions of this added project somewhere.

Let me know where I should capture it (for now I added it to issue #20). But here is my thinking:

  1. We need to have a set of tests that contributors can run to verify that their PRs have desired impact on perf, i.e. they would run the tests on master, then on the branch with changes, and compare the results.
  2. This project is not to do comparative benchmarking between different ML frameworks.


PredictionModel<IrisData, IrisPrediction> model = pipeline.Train<IrisData, IrisPrediction>();

IrisPrediction prediction = model.Predict(new IrisData()
Copy link
Contributor

@glebuk glebuk May 9, 2018

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Predict [](start = 46, length = 7)

Best to separate into two separate tests: 1. for training another for interence/predictions.
Those two are measured independently and are tuned independently #Closed

Copy link
Member Author

@KrzysztofCwalina KrzysztofCwalina May 9, 2018

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good point. I split training and prediction into two separate tests #Closed

PetalWidth = 5.1f,
});

prediction = model.Predict(new IrisData()
Copy link
Contributor

@glebuk glebuk May 9, 2018

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Predict [](start = 31, length = 7)

There are two dimensions on measuring predictions:

  1. Measiuring pipeline creation time
  2. Measuring each individual prediction once pipeline is established.
    Often the two differ by 100x. Without this separation the results are very misleading
    In high perf systems a great care must be taken to minimize initialization of predictors. #Closed

Copy link
Member Author

@KrzysztofCwalina KrzysztofCwalina May 9, 2018

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is it in addition to splitting training and prediction? I created two models (one cold and one hot) and ran the benchmark. There was no difference in the results. #Closed

Copy link
Contributor

@glebuk glebuk May 10, 2018

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Quick question - can you compare times of running predictions with an input with 1,2 and 5 examples per function call?. Are they linear (1,2 5x)? I believe the pipelin is iniialized once per enumeration
If the times for tests above are nearly equal, then it means that most time is taken with pipeline initialization and not prediction.
In this case you need to call predict against an IEnumerable
In this case., because the model is tiny, the effect might be small, however, if models are bug, the model initialization might dominate.


In reply to: 187092994 [](ancestors = 187092994)

Copy link
Member Author

@KrzysztofCwalina KrzysztofCwalina May 10, 2018

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

They scale nicely, as if the pipeline (model) was fully initialized when it's created.

BenchmarkDotNet=v0.10.14, OS=Windows 10.0.16299.431 (1709/FallCreatorsUpdate/Redstone3)
Intel Core i7-6700 CPU 3.40GHz (Skylake), 1 CPU, 8 logical and 4 physical cores
Frequency=3328124 Hz, Resolution=300.4696 ns, Timer=TSC
.NET Core SDK=2.1.300-rc1-008662
  [Host] : .NET Core 2.0.7 (CoreCLR 4.6.26328.01, CoreFX 4.6.26403.03), 64bit RyuJIT

Toolchain=InProcessToolchain  LaunchCount=1  TargetCount=3  
WarmupCount=3  
Method NumberOfPredictions Mean Error StdDev AccuracyMacro Gen 0 Gen 1 Gen 2 Allocated
PredictIris 1 1.011 ms 0.4383 ms 0.0248 ms 0.98 27.3438 13.6719 1.9531 88.15 KB
PredictIris 2 2.145 ms 3.0685 ms 0.1734 ms 0.98 54.6875 27.3438 3.9063 176.31 KB
PredictIris 5 5.326 ms 8.2691 ms 0.4672 ms 0.98 132.8125 70.3125 15.6250 440.45 KB

Copy link
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool. Is it possible to add a test that takes an enumeration of 1, 2, and 5 data points for each call of Predict? I"d expect it to be a lot faster.


In reply to: 187421073 [](ancestors = 187421073)

Copy link
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Need to run tests with several examples for a single Predict call, not several calls to predict. They are not the same.


In reply to: 187246794 [](ancestors = 187246794,187092994)


namespace Microsoft.ML.Benchmarks
{
public class TrainPredictionBench
Copy link
Contributor

@glebuk glebuk May 9, 2018

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

TrainPredictionBench [](start = 17, length = 20)

Split benchmarks by task (regression, binary, etc). You already have classification metrics as a field. Perhaps have a base class and then subclasses for each task. #Closed

Copy link
Member Author

@KrzysztofCwalina KrzysztofCwalina May 9, 2018

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

changed the test to StochasticDualCoordinateAscentClassifierBench #Closed

Copy link
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should the file name be updated too?


In reply to: 187092340 [](ancestors = 187092340)

Copy link
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you please update the file name to match the class name?


In reply to: 187243814 [](ancestors = 187243814,187092340)

{
public class TrainPredictionBench
{
internal static ClassificationMetrics s_metrics;
Copy link
Contributor

@glebuk glebuk May 9, 2018

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nternal static ClassificationMetrics s_metrics; [](start = 9, length = 47)

Would also be awesome to figure out how to minimize the amount of code per dataset. so that you can just add a dataset and then not have to change much else
Also, it should be good to be able to quickly and easily substitute different learners.
Basically the dimensions for tests are:
Training:
Task:
Learner
Hyper-parameters{ detault, sweep}
Dataset
Subset by size (like 10%, 50%, 100% of rows)
Inference:
Pipeline load time
Time per prediction #Pending

Copy link
Member Author

@KrzysztofCwalina KrzysztofCwalina May 9, 2018

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, this is a good idea, but I think such generalizations are best done when actually trying to develop the second, third, etc. test. i.e. I will try to do it when I am adding more tests. #Closed

@KrzysztofCwalina
Copy link
Member Author

@dotnet-bot test OSX10.13 Debug


namespace Microsoft.ML.Benchmarks
{
class Program
Copy link
Contributor

@glebuk glebuk May 10, 2018

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Program [](start = 10, length = 7)

Why is this an EXE and not a test?
I wonder would it be easier if we just have such tests as regular unit tests? from a dev perpective, we have nice tools to run and compare results for tests. #Pending

Copy link
Member Author

@KrzysztofCwalina KrzysztofCwalina May 10, 2018

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

BenchmarkDotNet (BDN) test are exes. I am pretty sure they cannot be dlls that are run as part of unit tests. I think @adamsitnik is working on infrastructure that will let us run the tests in the outer loop (daily runs) #Resolved

SepalWidth = 1.6f,
PetalLength = 0.2f,
PetalWidth = 5.1f,
});
Copy link
Contributor

@glebuk glebuk May 14, 2018

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: Probably better to create those separately as static enumeration of test data. that way you don't measure the object creation time
Then you can easily call different methods on differnt length of prediction enums or single objects by calling TestData.Take(5); or TestData.First(); #Closed

Copy link
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done, though First and Take allocate way more :-)

Copy link
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

sounds good


In reply to: 188046953 [](ancestors = 188046953)

Copy link
Contributor

@glebuk glebuk left a comment

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

Copy link
Contributor

@glebuk glebuk left a comment

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@eerhardt
Copy link
Member

(nit) ToolsVersion="15.0" xmlns="[http://schemas.microsoft.com/developer/msbuild/2003](http://schemas.microsoft.com/developer/msbuild/2003)" are not used, and unnecessary.


Refers to: test/Microsoft.ML.Benchmarks/Microsoft.ML.Benchmarks.csproj:1 in 0a281b2. [](commit_id = 0a281b2, deletion_comment = False)

<OutputType>Exe</OutputType>
<LangVersion>7.2</LangVersion>
<TargetFramework>netcoreapp2.0</TargetFramework>
<StartupObject>Microsoft.ML.Benchmarks.Program</StartupObject>
Copy link
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What is this used for? I've never seen this in new SDK-style projects.

Copy link
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

When I remove it, it tells me there is more than one entry point and so the one real one must be specified. Possibly one of the BDN dlls has a Main method? @adamsitnik?

}
}

private static PredictionModel<IrisData, IrisPrediction> TrainCore()
Copy link
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this method a duplicate of TrainIris()?

Copy link
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, they used to be different, but now after all the tweaks they become identical. I will remove one of them.

Copy link
Member

@eerhardt eerhardt left a comment

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@KrzysztofCwalina KrzysztofCwalina merged commit 83f9bac into dotnet:master May 15, 2018
eerhardt pushed a commit to eerhardt/machinelearning that referenced this pull request Jul 27, 2018
* Benchmark

* Changed to .NET Core app

* Added Accuracy Reporting

* fixed build

* Feedback from Gleb

* Added batch prediction tests

* Resolved conflicts the sln file

* Renamed the new file to match type name

* Removed duplicated method
Dmitry-A pushed a commit to Dmitry-A/machinelearning that referenced this pull request Apr 12, 2019
2) Fixing up samples to reflect it
Dmitry-A added a commit that referenced this pull request Apr 13, 2019
…ature branch (#3324)

* Initial commit

* ci test build

* forgot to save this one file

* Debug-Intrinsics isn't a valid config, trying windows-x64

* disabled tests for now

* disable tests attempt 2

* initial code push, no history, test project not in the build so is the internal client

* battling with warn as err

* test build

* test change

* make params for MLContext data extensions match ML.NET default names and values; update gitignore; nit rev for Benchmarking.cs (#5)

* Create README.md (#2)

* API folder changes (#6)

* comment out fast forest trainer, per discussion on ML.NET open issue #1983, for now, to run E2E w/o exceptions (#7)

* Make validation data param mandatory; remove GetFirstPipeline sample (#10)

* Make validation data param mandatory; remove GetFirstPipeline sample

* remove deprecated todo

* Create ISSUE_TEMPLATE.md & PULL_REQUEST_TEMPLATE.md (#12)

* Create ISSUE_TEMPLATE.md

* Create PULL_REQUEST_TEMPLATE.md

* NestedObject For pipeline (#14)

* add estimator extensions / catalog; add conversion from external to internal pipeline; transform clean-up; add back in test proj and fix build; refactor trainer ext name mappings (#15)

* Make validation data param mandatory; remove GetFirstPipeline sample

* remove deprecated todo

* add estimator extensions / catalog; add ability to go from external to internal pipeline; a lot of transform clean-up; add back in test proj and get it building; refactor trainer ext name mappings

* corrected the typo in readme (#16)

* make GetNextPipeline API w/ public Pipeline method on PipelineSuggester; write GetNextPipeline API test; fix public Pipeline object serialization; fix header inferencing bug; write test utils for fetching datasets (#18)

* get next pipeline API rev -- refactor API to consume column dimensions, purpose, type, and name instead of available trainers & transforms (#19)

* mark get next pipeline test as ignore for now (#20)

* fix dataview take util bug, add dataview skip util, add some UTs to increase code coverage (#21)

* fix dataview take util bug, add dataview skip util, add some UTs to increase code coverage

* add accuracy threshold on AutoFit test

* add null check to best pipeline on autofit result

* unit test additions (including user input validation testing); dead code removal for code coverage (including KDO & associated utils); misc fixes & revs (#22)

* add trainer extension tests, & misc fixes (#23)

* add estimator extension tests (#24)

* add conversions tests (#25)

* fix multiclass runs & add multiclass autofit UT (#27)

* add basic autofit regression test (#28)

* fix categorical transform bug (sometimes categorical features weren't concatenated to final features); add UT transforms; add PipelineNode equality & tests to serve as AutoML testing infra

* add example to readme (#26)

* add lightgbm args as nested properties (#33)

* fix bug where if one pipeline hyperparam optimization converges, run terminates (#36)

* add open-source headers to files; other nit clean-ups along the way (#35)

* Ungroup Columns in Column Inference (#40)

* Added sequential grouping of columns

* added ungrouping of column option

* reverted the file

* Misc fixes (#39)

* misc fixes -- fix bug where SMAC returning already-seen values; fix param encoding return bug in pipeline object model; nit clean-up AutoFit; return in pipeline suggester when sweeper has no next proposal; null ref fix in public object model pipeline suggester

* fix in BuildPipelineNodePropsLightGbm test, fix / use correct 'newTrainer' variable in PipelneSuggester

* SMAC perf improvement

* Removing the nuget.config and have build.props mention the nuget package sources. (#38)

* Added sequential grouping of columns

* removed nuget.config and have only props mentions the nuget sources

* reverted the file

* transform inferencing concat / ignore fixes (#41)

* make pipeline object model & other public classes internal (#43)

* handle SMAC exception when fewer trees were trained than requested (#44)

* Throw error on incorrect Label name in InferColumns API (#47)

* Added sequential grouping of columns

* reverted the file

* addded infer columns label name checking

* added column detection error

* removed unsed usings

* added quotes

* replace Where with Any clause

* replace Where with Any clause

* Set Nullable Auto params to null values (#50)

* Added sequential grouping of columns

* reverted the file

* added auto params as null

* change to the update fields method

* First public api propsal (#52)

* Includes following
1) Final proposal for 0.1 public API surface
2) Prefeaturization
3) Splitting train data into train and validate when validation data is null
4) Providing end to end samples one each for regression, binaryclassification and multiclass classification

* Incorporating code review feedbacks

* Revert "Set Nullable Auto params to null values" (#53)

* Revert "First public api propsal (#52)"

This reverts commit e4a64cf.

* Revert "Set Nullable Auto params to null values (#50)"

This reverts commit 41c663c.

* AutoFit return type is now an IEnumerable (#55)

AutoFit returns is now an IEnumerable - this enables many good things

Implementing variety of early stopping criteria (See sample)
Early discard of models that are no good. This improves memory usage efficiency. (See sample)
No need to implement a callback to get results back
Getting best score is now outside of API implementation. It is a simple math function to compare scores (See sample).

Also templatized the return type for better type safety through out the code.

* misc fixes & test additions, towards 0.1 release (#56)

* Enable UnitTests on build server (#57)

* 1) Making trainer name public (#62)

2) Fixing up samples to reflect it

*  Initial version of CLI tool for mlnet (#61)

* added global tool initial project

* removed unneccesary files, renamed files

* refactoring and added base abstract classes for trainer generator

* removed unused class

* Added classes for transforms

* added transform generate dummy classes

* more refactoring, added first transform

* more refactoring and added classes

* changed the project structure

* restructing added options class

* sln changes

* refactored options to different class:

* added more logic for code generation of class

* misc changes

* reverted file

* added commandline api package

* reverted sample

* added new command line api parser

* added normalization of column names

* Added command defaults and error message

* implementation of all trainers

* changed auto to null

* added all transform generators

* added error handling when args is empty and minor changes due to change in AutoML api names

* changed the name of param

* added new command line options and restructuring code

* renamed proj file and added solution

* Added code to generate usings, Fixed few bugs in the code

* added validation to the command line options

* changed project name

* Bug fixes due to API change in AutoML

* changed directory structure

* added test framework and basic tests

* added more tests

* added improvements to template and error handling

* renamed the estimator name

* fixed test case

* added comments

* added headers

* changed namespace and removed unneccesary properties from project

* Revert "changed namespace and removed unneccesary properties from project"

This reverts commit 9edae033e9845e910f663f296e168f1182b84f5f.

* fixed test cases and renamed namespaces

* cleaned up proj file

* added folder structure

* added symbols/tokens for strings

* added more tests

* review comments

* modified test cases

* review comments

* change in the exception message

* normalized line endings

* made method private static

* simplified range building /optimization

* minor fix

* added header

* added static methods in command where necessary

* nit picks

*  made few methods static

* review comments

* nitpick

* remove line pragmas

* fix test case

* Use better AutiFit overload and ignore Multiclass (#64)

* Upgrading CLI to produce ML.NET V.10 APIs and bunch of Refactoring tasks (#65)

* Added sequential grouping of columns

* reverted the file

* upgrade to v .10 and refactoring

* added null check

* fixed unit tests

* review comments

* removed the settings change

* added regions

* fixed unit tests

* Upgrade ML.NET package to 0.10.0 (#70)

* Change in template to accomodate new API of TextLoader (#72)

* Added sequential grouping of columns

* reverted the file

* changed to new API of Text Loader

* changed signature

* added params for taking additional settings

* changes to codegen params

* refactoring of templates and fixing errors

* Enable gated check for mlnet.tests (#79)

* Added sequential grouping of columns

* reverted the file

* changed to new API of Text Loader

* changed signature

* added params for taking additional settings

* changes to codegen params

* refactoring of templates and fixing errors

* added run-tests.proj and referred it in build.proj

* CLI tool - make validation dataset optional and support for crossvalidation in generated code (#83)

* Added sequential grouping of columns

* reverted the file

* bug fixes, more logic to templates to support cross-validate

* formatting and fix type in consolehelper

* Added logic in templates

* revert settings

* benchmarking related changes (#63)

* Create test.txt

* Create test.txt

* changes needed for benchmarking

* forgot one file

* merge conflict fix

* fix build break

* back out my version of the fix for Label column issue and fix the original fix

* bogus file removal

* undo SuggestedPipeline change

* remove labelCol from pipeline suggester

* fix build break

* fix fast forest learner (don't sweep over learning rate) (#88)

* Made changes to Have non-calibrated scoring for binary classifiers (#86)

* Added sequential grouping of columns

* reverted the file

* added calibration workaround

* removed print probability

* reverted settings

* rev ColumnInference API: can take label index; rev output object types; add tests (#89)

* rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (#99)

* Create test.txt

* Create test.txt

* changes needed for benchmarking

* forgot one file

* merge conflict fix

* fix build break

* back out my version of the fix for Label column issue and fix the original fix

* bogus file removal

* undo SuggestedPipeline change

* remove labelCol from pipeline suggester

* fix build break

* rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline)

* publish nuget (#101)

* use dotnet-internal-temp agent for internal build

* use dotnet-internal feed

* Fix Codegen for columnConvert and ValueToKeyMapping transform and add individual transform tests (#95)

* Added sequential grouping of columns

* reverted the file

* fix usings for type convert

* added transforms tests

* review comments

* When generating usings choose only distinct usings directives (#94)

* Added sequential grouping of columns

* reverted the file

* Added code to have unique strings

* refactoring

* minor fix

* minor fix

* Autofit overloads + cancellation + progress callbacks

1) Introduce AutoFit overloads (basic and advanced)
2) AutoFit Cancellation
3) AutoFit progress callbacks

* Default the kfolds to value 5 in CLI generated code (#115)

* Added sequential grouping of columns

* reverted the file

* Set up CI with Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* remove file

* added kfold param and defaulted to value

* changed type

* added for regression

* Remove extra ; from generated code (#114)

* Added sequential grouping of columns

* reverted the file

* Set up CI with Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* removed extra ; from generated code

* removed file

* fix unit tests

* TimeoutInSeconds (#116)

Specifying timeout in seconds instead of minutes

* Added more command line args implementation to CLI tool and refactoring (#110)

* Added sequential grouping of columns

* reverted the file

* Set up CI with Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* added git status

* reverted change

* added codegen options and refactoring

* minor fixes'

* renamed params, minor refactoring

* added tests for commandline and refactoring

* removed file

* added back the test case

* minor fixes

* Update src/mlnet.Test/CommandLineTests.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* review comments

*  capitalize the first character

* changed the name of test case

* remove unused directives

* Fail gracefully if unable to instantiate data view with swept parameters (#125)

* gracefully fail if fail to parse a datai

* rev

* validate AutoFit 'Features' column must be of type R4 (#132)

* Samples: exceptions / nits (#124)

* Logging support in CLI + Implementation of cmd args [--name,--output,--verbosity] (#121)

* addded logging and helper methods

* fixing code after merge

* added resx files, added logger framework, added logging messages

* added new options

* added spacing

* minor fixes

* change command description

* rename option, add headers, include new param in test

* formatted

* build fix

*  changed option name

* Added NlogConfig file

* added back config package

* fix tests

* added correct validation check (#137)

* Use CreateTextLoader<T>(..)  instead of CreateTextLoader(..) (#138)

* added support to loaddata by class in the generated code

* fix tests

* changed CreateTextLoader to ReadFromTextFile method. (#140)

* changed textloader to readfromtextfile method

* formatting

* exception fixes (#136)

* infer purpose of hidden columns as 'ignore' (#142)

* Added approval tests and bunch of refactoring of code and normalizing namespaces (#148)

* changed textloader to readfromtextfile method

* formatting

* added approval tests and refactoring of code

* removed few comments

* API 2.0 skeleton (#149)

Incorporating API review feedback

* The CV code should come before the training when there is no test dataset in generated code (#151)

* reorder cv code

* build fix

* fixed structure

* Format the generated code + bunch of misc tasks (#152)

* added formatting and minor changes for reordering cv

* fixing the template

* minor changes

* formatting changes

* fixed approval test

* removed unused nuget

* added missing value replacing

* added test for new transform

* fix test

* Update src/mlnet/Templates/Console/MLCodeGen.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* Sanitize the column names in CLI (#162)

* added sanitization layer in CLI

* fix test

* changed exception.StackTrace to exception.ToString()

* fix package name (#168)

* Rev public API (#163)

* Rename TransformGeneratorBase .cs to TransformGeneratorBase.cs (#153)

* Fix minor version for the repository + remove Nlog config package (#171)

*  changed the minor version

* removed the nlog config package

* Added new test to columninfo and fixing up API (#178)

* Make optimizing metric customizable and add trainer whitelist functionality (#172)

* API rev (#181)

* propagate root MLContext thru AutoML (instead of creating our own) (#182)

* Enabling new command line args (#183)

* fix package name

* initial commit

* added more commandline args

* fixed tests

* added headers

* fix tests

* fix test

* rename 'AutoFitter' to 'Experiment' (#169)

* added tests (#187)

* rev InferColumns to accept ColumnInfo input param (#186)

* Implement argument --has-header and change usage of dataset (#194)

* added has header and fixed dataset and train dataset

* fix tests

* removed dummy command (#195)

* Fix bug for regression and sanitize input label from user (#198)

* removed dummy command

* sanitize label and fix template

* fix tests

* Do not generate code concatenating columns when the dataset has a single feature column (#191)

* Include some missed logging in the generated code.  (#199)

* added logging messages for generated code

* added log messages

* deleted file

* cleaning up proj files (#185)

* removed platform target

* removed platform target

* Some spaces and extra lines + bug in output path  (#204)

* nit picks

* nit picks

* fix test

* accept label from user input and provide in generated code (#205)

* Rev handling of weight / label columns (#203)

* migrate to private ML.NET nuget for latest bug fixes (#131)

* fix multiclass with nonstandard label (#207)

* Multiclass nondefault label test (#208)

* printing escaped chars + bug (#212)

* delete unused internal samples (#211)

* fix SMAC bug that causes multiclass sample to infinite loop (#209)

* Rev user input validation for new API (#210)

* added console message for exit and nit picks (#215)

* exit when exception encountered (#216)

* Seal API classes (and make EnableCaching internal) (#217)

* Suggested sample nits (feel free to ask for any of these to be reverted) (#219)

* User input column type validation (#218)

* upgrade commandline and renaming (#221)

* upgrade commandline and renaming

* renaming fields

* Make build.sh, init-tools.sh, & run.sh executable on OSX/Linux (#225)

*  CLI argument descriptions updated (#224)

* CLI argument descriptions updated

* No version in .csproj

* added flag to disable training code (#227)

* Exit if perfect model produced (#220)

* removed header (#228)

* removed header

* added auto generated header

* removed console read key (#229)

* Fix model path in generated file (#230)

* removed console read key

* fix model path

* fix test

* reorder samples (#231)

* remove rule that infers column purpose as categorical if # of distinct values is < 100 (#233)

* Null reference exception fix for finding best model when some runs have failed (#239)

* samples fixes (#238)

* fix for defaulting Averaged Perceptron # of iterations to 10 (#237)

* Bug bash feedback Feb 27. API changes and sample changes (#240)

* Bug bash feedback Feb 27. 
API changes 
Sample changes
Exception fix

* Samples / API rev from 2/27 bug bash feedback (#242)

* changed the directory structure for generated project (#243)

* changed the directory structure for generated project

* changed test

* upgraded commandline package

* Fix test file locations on OSX (#235)

* fix test file locations on OSX

* changing to Path.Combine()

* Additional Path.Combine()

* Remove ConsoleCodeGeneratorTests.GeneratedTrainCodeTest.received.txt

* Additional Path.Combine()

* add back in double comparison fix

* remove metrics agent NaN returns

* test fix

* test format fix

* mock out path

Thanks to @daholste for additional fixes!

* upgrade to latest ML.NET public surface (#246)

* Upgrade to ML.NET 0.11 (#247)

* initial changes

* fix lightgbm

* changed normalize method

* added tests

* fix tests

* fix test

* Private preview final API changes (#250)

* .NET framework design guidelines applied to public surface
* WhitelistedTrainers -> Trainers

* Add estimator to public API iteration result (#248)

* LightGBM pipeline serialization fix (#251)

* Change order that we search for TextLoader's parameters (#256)

* CLI IFileInfo null exception fix (#254)

* Averaged Perceptron pipeline serialization fix (#257)

* Upgrade command-line-api and default folder name change (#258)

* change in defautl folderName

* upgrade command line

* Update src/mlnet/Program.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* eliminate IFileInfo from CLI (#260)

* Rev samples towards private preview; ignored columns fix (#259)

* remove unused methods in consolehelper and nit picks in generated code (#261)

* nit picks

* change in console helper

* fix tests

* add space

* fix tests

* added nuget sources in generated csproj (#262)

* added nuget sources in csproj

* changed the structure in generated code

* space

* upgrade to mlnet 0.11 (#263)

* Formatting CLI metrics (#264)

Ensures space between printed metrics (also model counter). Right aligned metrics. Extended AUC to four digits.

* Add implementation of non -ova multi class trainers code gen (#267)

* added non ova multi class learners

* added tests

* test cases

* Add caching (#249)

* AdvancedExperimentSettings sample nits (#265)

* Add sampling key column (#268)

* Initial work for multi-class classification support for CLI (#226)

* Initial work for multi-class classification support for CLI

* String updates

* more strings

* Whitelist non-OVA multi-class learners

* Refactor the orchestration of AutoML calls (#272)

* Do not auto-group columns with suggested purpose = 'Ignore' (#273)

* Fix: during type inferencing, parse whitespace strings as NaN (#271)

* Printing additional metrics in CLI for binary classification (#274)

* Printing additional metrics in CLI for binary classification

* Update src/mlnet/Utilities/ConsolePrinter.cs

* Add API option to store models on disk (instead of in memory); fix IEstimator memory leak (#269)

* Print failed iterations in CLI (#275)

* change the type to float from double (#277)

* cache arg implementation in CLI (#280)

* cache implementation

* corrected the null case

* added tests for all cases

* Remove duplicate value-to-key mapping transform for multiclass string labels (#283)

* Add post-trainer transform SDK infra; add KeyToValueMapping transform to CLI; fix: for generated multiclass models, convert predicted label from key to original label column type (#286)

* Implement ignore columns command line arg (#290)

* normalize line endings

* added --ignore-columns

* null checks

* unit tests

* Print winning iteration and runtime in CLI (#288)

* Print best metric and runtime

* Print best metric and runtime

* Line endings in AutoMLEngine.cs

* Rename time column to duration to match Python SDK

* Revert to MicroAccuracy and MacroAccuracy spellings

* Revert spelling of BinaryClassificationMetricsAgent to BinaryMetricsAgent to reduce merge conflicts

* Revert spelling of MulticlassMetricsAgent to MultiMetricsAgent to reduce merge conflicts

* missed some files

* Fix merge conflict

* Update AutoMLEngine.cs

* Add MacOS & Linux to CI; MacOS & Linux test fixes (#293)

* MicroAccuracy as default for multi-class (#295)

Change default optimization metric for multi-class classification to MicroAccuracy (accuracy). Previously it was set to MacroAccuracy.

* Null exception for ignorecolumns in CLI (#294)

* Null exception for ignorecolumns in CLI

* Check if ignore-columns array has values (as the default is now a empty array)

* Emit caching flag in pipeline object model. (Includes SuggestedPipelineBuilder refactor & debug string fixes / refactor) (#296)

* removed sln (#297)

* Caching enabling in code gen part -2 (#298)

* add

* added caching codegen

* support comma separated values for --ignore-columns (#300)

* default initialization for ignore columns (#302)

* default initialization

* adde null check

* Codegen for multiclass non-ova (#303)

* changes to template

* multicalss codegen

* test cases

* fix test cases

* Generated Project new structure. (#305)

* added new templates

* writing files to disck

* change path

* added new templates

* misisng braces

* fix bugs

* format code

* added util methods for solution file creation and addition of projects to it

* added extra packages to project files

* new tests

* added correct path for sln

* build fix

* fix build

* include using system in prediction class (#307)

* added using

* fix test

* Random number generator is not thread safe (#310)

* Random number generator is not thread safe

* Another local random generator

* Missed a few references

* Referncing AutoMlUtils.random instead of a local RNG

* More refs to mail RNG; remove Float as per #1669

* Missed Random.cs

* Fix multiclass code gen (#314)

* compile error in codegen

* removes scores printing

* fix bugs

* fix test

* Fix compile error in codegen project (#319)

* removed redundant code

* fix test case

* Rev OVA pipeline node SDK output: wrap binary trainers as children inside parent OVA node (#317)

* Ova Multi class codegen support (#321)

* dummy

* multiova implementation

* fix tests

* remove inclusion list

* fix tests and console helper

* Rev run result trainer name for OVA: output different trainer name for each OVA + binary learner combination (#322)

* Rev run result trainer name for Ova: output different trainer name for each Ova + binary learner combination

* test fixes

* Console helper bug in generated code for multiclass (#323)

* fix

* fix test

* looping perlogclass

* fix test

* Initial version of Progress bar impl and CLI UI experience (#325)

* progressbar

* added progressbar and refactoring

* reverted

* revert sign assembly

* added headers and removed exception rethrow

* Setting model directory to temp directory (#327)

* Suggested changes to progress bar (#335)

* progressbar

* added progressbar and refactoring

* reverted

* revert sign assembly

* added headers and removed exception rethrow

* bug fixes and updates to UI

* added friendly name printing for metric

* formatting

* Rev Samples (#334)

* Telemetry2 (#333)

* Create test.txt

* Create test.txt

* changes needed for benchmarking

* forgot one file

* merge conflict fix

* fix build break

* back out my version of the fix for Label column issue and fix the original fix

* bogus file removal

* undo SuggestedPipeline change

* remove labelCol from pipeline suggester

* fix build break

* rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline)

* tweak queue in vsts-ci.yml

* CLI telemetry implementation

* Telemetry implementation

* delete unnecessary file and change file size bucket to actually log log2 instead of nearest ceil value

* add headers, remove comments

* one more header missing

* Fix progress bar in linux/osx (#336)

* progressbar

* added progressbar and refactoring

* reverted

* revert sign assembly

* added headers and removed exception rethrow

* bug fixes and updates to UI

* added friendly name printing for metric

* formatting

* change from task to thread

* Update src/mlnet/CodeGenerator/CodeGenerationHelper.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* Mem leak fix (#328)

* Create test.txt

* Create test.txt

* changes needed for benchmarking

* forgot one file

* merge conflict fix

* fix build break

* back out my version of the fix for Label column issue and fix the original fix

* bogus file removal

* undo SuggestedPipeline change

* remove labelCol from pipeline suggester

* fix build break

* rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline)

* tweak queue in vsts-ci.yml

* there is still investigation to be done but this fix works and solves memory leak problems

* minor refactor

* Upgrade ML.NET package (#343)

* Add cross-validation (CV), and auto-CV for small datasets; push common API experiment methods into base class (#287)

* restore old yml for internal pipeline so we can publish nuget again to devdiv stream (#344)

* Polishing the CLI UI part-1 (#338)

* formatting of pbar message

* Polishing the UI

* optimization

* rename variable

* Update src/mlnet/AutoML/AutoMLEngine.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* Update src/mlnet/CodeGenerator/CodeGenerationHelper.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* new message

* changed hhtp to https

* added iteration num + 1

* change string name and add color to artifacts

* change the message

* build errors

* added null checks

* added exception messsages to log file

* added exception messsages to log file

* CLI ML.NET version upgrade (#345)

* Sample revs; ColumnInformation property name revs; pre-featurizer fixes (#346)

* CLI -- consume logs from AutoML SDK (#349)

* Rename RunDetails --> RunDetail (#350)

* command line api upgrade and progress bar rendering bug (#366)

* added fix for all platforms progress bar

* upgrade nuget

* removed args from writeline

* change in the version (#368)

* fix few bugs in progressbar and verbosity (#374)

* fix few bugs in progressbar and verbosity

* removed unused name space

* Fix for folders with space in it while generating project (#376)

* support for folders with spaces

* added support for paths with space

* revert file

* change name of var

* remove spaces

* SMAC fix for minimizing metrics (#363)

* Formatting Regression metrics and progress bar display days. (#379)

* added progress bar day display and fix regression metrics

* fix formatting

* added total time

* formatted total time

* change command name and add pbar message (#380)

* change command name and add pbar message

* fix tests

* added aliases

* duplicate alias

* added another alias for task

* UI missing features (#382)

* added formatting changes

* added accuracy specifically

* downgrade the codepages (#384)

* Change in project structure (#385)

* initial changes

* Change in project structure

* correcting test

* change variable name

* fix tests

* fix tests

* fix more tests

* fix codegen errors

* adde log file message

* changed name of args

* change variable names

* fix test

* FileSizeBuckets in correct units (#387)

* Minor telemetry change to log in correct units and make our life easier in the future

* Use Ceiling instead of Round

* changed order (#388)

* prep work to transfer to ml.net (#389)

* move test projects to top level test subdir

* rename some projects to make naming consistent and make it build again

* fix test project refs

* Add AutoML components to build, fix issues related to that so it builds
Dmitry-A pushed a commit to Dmitry-A/machinelearning that referenced this pull request Aug 22, 2019
2) Fixing up samples to reflect it
harishsk added a commit that referenced this pull request Sep 6, 2019
* Fixed build errors resulting from upgrade to VS2019 compilers

* Added additional message describing the previous fix

* Syncing upstream fork (#10)

* Throw error on incorrect Label name in InferColumns API (#47)

* Added sequential grouping of columns

* reverted the file

* addded infer columns label name checking

* added column detection error

* removed unsed usings

* added quotes

* replace Where with Any clause

* replace Where with Any clause

* Set Nullable Auto params to null values (#50)

* Added sequential grouping of columns

* reverted the file

* added auto params as null

* change to the update fields method

* First public api propsal (#52)

* Includes following
1) Final proposal for 0.1 public API surface
2) Prefeaturization
3) Splitting train data into train and validate when validation data is null
4) Providing end to end samples one each for regression, binaryclassification and multiclass classification

* Incorporating code review feedbacks

* Revert "Set Nullable Auto params to null values" (#53)

* Revert "First public api propsal (#52)"

This reverts commit e4a64cf4aeab13ee9e5bf0efe242da3270241bd7.

* Revert "Set Nullable Auto params to null values (#50)"

This reverts commit 41c663cd14247d44022f40cf2dce5977dbab282d.

* AutoFit return type is now an IEnumerable (#55)

AutoFit returns is now an IEnumerable - this enables many good things

Implementing variety of early stopping criteria (See sample)
Early discard of models that are no good. This improves memory usage efficiency. (See sample)
No need to implement a callback to get results back
Getting best score is now outside of API implementation. It is a simple math function to compare scores (See sample).

Also templatized the return type for better type safety through out the code.

* misc fixes & test additions, towards 0.1 release (#56)

* Enable UnitTests on build server (#57)

* 1) Making trainer name public (#62)

2) Fixing up samples to reflect it

*  Initial version of CLI tool for mlnet (#61)

* added global tool initial project

* removed unneccesary files, renamed files

* refactoring and added base abstract classes for trainer generator

* removed unused class

* Added classes for transforms

* added transform generate dummy classes

* more refactoring, added first transform

* more refactoring and added classes

* changed the project structure

* restructing added options class

* sln changes

* refactored options to different class:

* added more logic for code generation of class

* misc changes

* reverted file

* added commandline api package

* reverted sample

* added new command line api parser

* added normalization of column names

* Added command defaults and error message

* implementation of all trainers

* changed auto to null

* added all transform generators

* added error handling when args is empty and minor changes due to change in AutoML api names

* changed the name of param

* added new command line options and restructuring code

* renamed proj file and added solution

* Added code to generate usings, Fixed few bugs in the code

* added validation to the command line options

* changed project name

* Bug fixes due to API change in AutoML

* changed directory structure

* added test framework and basic tests

* added more tests

* added improvements to template and error handling

* renamed the estimator name

* fixed test case

* added comments

* added headers

* changed namespace and removed unneccesary properties from project

* Revert "changed namespace and removed unneccesary properties from project"

This reverts commit 9edae033e9845e910f663f296e168f1182b84f5f.

* fixed test cases and renamed namespaces

* cleaned up proj file

* added folder structure

* added symbols/tokens for strings

* added more tests

* review comments

* modified test cases

* review comments

* change in the exception message

* normalized line endings

* made method private static

* simplified range building /optimization

* minor fix

* added header

* added static methods in command where necessary

* nit picks

*  made few methods static

* review comments

* nitpick

* remove line pragmas

* fix test case

* Use better AutiFit overload and ignore Multiclass (#64)

* Upgrading CLI to produce ML.NET V.10 APIs and bunch of Refactoring tasks (#65)

* Added sequential grouping of columns

* reverted the file

* upgrade to v .10 and refactoring

* added null check

* fixed unit tests

* review comments

* removed the settings change

* added regions

* fixed unit tests

* Upgrade ML.NET package to 0.10.0 (#70)

* Change in template to accomodate new API of TextLoader (#72)

* Added sequential grouping of columns

* reverted the file

* changed to new API of Text Loader

* changed signature

* added params for taking additional settings

* changes to codegen params

* refactoring of templates and fixing errors

* Enable gated check for mlnet.tests (#79)

* Added sequential grouping of columns

* reverted the file

* changed to new API of Text Loader

* changed signature

* added params for taking additional settings

* changes to codegen params

* refactoring of templates and fixing errors

* added run-tests.proj and referred it in build.proj

* CLI tool - make validation dataset optional and support for crossvalidation in generated code (#83)

* Added sequential grouping of columns

* reverted the file

* bug fixes, more logic to templates to support cross-validate

* formatting and fix type in consolehelper

* Added logic in templates

* revert settings

* benchmarking related changes (#63)

* Create test.txt

* Create test.txt

* changes needed for benchmarking

* forgot one file

* merge conflict fix

* fix build break

* back out my version of the fix for Label column issue and fix the original fix

* bogus file removal

* undo SuggestedPipeline change

* remove labelCol from pipeline suggester

* fix build break

* fix fast forest learner (don't sweep over learning rate) (#88)

* Made changes to Have non-calibrated scoring for binary classifiers (#86)

* Added sequential grouping of columns

* reverted the file

* added calibration workaround

* removed print probability

* reverted settings

* rev ColumnInference API: can take label index; rev output object types; add tests (#89)

* rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (#99)

* Create test.txt

* Create test.txt

* changes needed for benchmarking

* forgot one file

* merge conflict fix

* fix build break

* back out my version of the fix for Label column issue and fix the original fix

* bogus file removal

* undo SuggestedPipeline change

* remove labelCol from pipeline suggester

* fix build break

* rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline)

* publish nuget (#101)

* use dotnet-internal-temp agent for internal build

* use dotnet-internal feed

* Fix Codegen for columnConvert and ValueToKeyMapping transform and add individual transform tests (#95)

* Added sequential grouping of columns

* reverted the file

* fix usings for type convert

* added transforms tests

* review comments

* When generating usings choose only distinct usings directives (#94)

* Added sequential grouping of columns

* reverted the file

* Added code to have unique strings

* refactoring

* minor fix

* minor fix

* Autofit overloads + cancellation + progress callbacks

1) Introduce AutoFit overloads (basic and advanced)
2) AutoFit Cancellation
3) AutoFit progress callbacks

* Default the kfolds to value 5 in CLI generated code (#115)

* Added sequential grouping of columns

* reverted the file

* Set up CI with Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* remove file

* added kfold param and defaulted to value

* changed type

* added for regression

* Remove extra ; from generated code (#114)

* Added sequential grouping of columns

* reverted the file

* Set up CI with Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* removed extra ; from generated code

* removed file

* fix unit tests

* TimeoutInSeconds (#116)

Specifying timeout in seconds instead of minutes

* Added more command line args implementation to CLI tool and refactoring (#110)

* Added sequential grouping of columns

* reverted the file

* Set up CI with Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* added git status

* reverted change

* added codegen options and refactoring

* minor fixes'

* renamed params, minor refactoring

* added tests for commandline and refactoring

* removed file

* added back the test case

* minor fixes

* Update src/mlnet.Test/CommandLineTests.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* review comments

*  capitalize the first character

* changed the name of test case

* remove unused directives

* Fail gracefully if unable to instantiate data view with swept parameters (#125)

* gracefully fail if fail to parse a datai

* rev

* validate AutoFit 'Features' column must be of type R4 (#132)

* Samples: exceptions / nits (#124)

* Logging support in CLI + Implementation of cmd args [--name,--output,--verbosity] (#121)

* addded logging and helper methods

* fixing code after merge

* added resx files, added logger framework, added logging messages

* added new options

* added spacing

* minor fixes

* change command description

* rename option, add headers, include new param in test

* formatted

* build fix

*  changed option name

* Added NlogConfig file

* added back config package

* fix tests

* added correct validation check (#137)

* Use CreateTextLoader<T>(..)  instead of CreateTextLoader(..) (#138)

* added support to loaddata by class in the generated code

* fix tests

* changed CreateTextLoader to ReadFromTextFile method. (#140)

* changed textloader to readfromtextfile method

* formatting

* exception fixes (#136)

* infer purpose of hidden columns as 'ignore' (#142)

* Added approval tests and bunch of refactoring of code and normalizing namespaces (#148)

* changed textloader to readfromtextfile method

* formatting

* added approval tests and refactoring of code

* removed few comments

* API 2.0 skeleton (#149)

Incorporating API review feedback

* The CV code should come before the training when there is no test dataset in generated code (#151)

* reorder cv code

* build fix

* fixed structure

* Format the generated code + bunch of misc tasks (#152)

* added formatting and minor changes for reordering cv

* fixing the template

* minor changes

* formatting changes

* fixed approval test

* removed unused nuget

* added missing value replacing

* added test for new transform

* fix test

* Update src/mlnet/Templates/Console/MLCodeGen.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* Sanitize the column names in CLI (#162)

* added sanitization layer in CLI

* fix test

* changed exception.StackTrace to exception.ToString()

* fix package name (#168)

* Rev public API (#163)

* Rename TransformGeneratorBase .cs to TransformGeneratorBase.cs (#153)

* Fix minor version for the repository + remove Nlog config package (#171)

*  changed the minor version

* removed the nlog config package

* Added new test to columninfo and fixing up API (#178)

* Make optimizing metric customizable and add trainer whitelist functionality (#172)

* API rev (#181)

* propagate root MLContext thru AutoML (instead of creating our own) (#182)

* Enabling new command line args (#183)

* fix package name

* initial commit

* added more commandline args

* fixed tests

* added headers

* fix tests

* fix test

* rename 'AutoFitter' to 'Experiment' (#169)

* added tests (#187)

* rev InferColumns to accept ColumnInfo input param (#186)

* Implement argument --has-header and change usage of dataset (#194)

* added has header and fixed dataset and train dataset

* fix tests

* removed dummy command (#195)

* Fix bug for regression and sanitize input label from user (#198)

* removed dummy command

* sanitize label and fix template

* fix tests

* Do not generate code concatenating columns when the dataset has a single feature column (#191)

* Include some missed logging in the generated code.  (#199)

* added logging messages for generated code

* added log messages

* deleted file

* cleaning up proj files (#185)

* removed platform target

* removed platform target

* Some spaces and extra lines + bug in output path  (#204)

* nit picks

* nit picks

* fix test

* accept label from user input and provide in generated code (#205)

* Rev handling of weight / label columns (#203)

* migrate to private ML.NET nuget for latest bug fixes (#131)

* fix multiclass with nonstandard label (#207)

* Multiclass nondefault label test (#208)

* printing escaped chars + bug (#212)

* delete unused internal samples (#211)

* fix SMAC bug that causes multiclass sample to infinite loop (#209)

* Rev user input validation for new API (#210)

* added console message for exit and nit picks (#215)

* exit when exception encountered (#216)

* Seal API classes (and make EnableCaching internal) (#217)

* Suggested sample nits (feel free to ask for any of these to be reverted) (#219)

* User input column type validation (#218)

* upgrade commandline and renaming (#221)

* upgrade commandline and renaming

* renaming fields

* Make build.sh, init-tools.sh, & run.sh executable on OSX/Linux (#225)

*  CLI argument descriptions updated (#224)

* CLI argument descriptions updated

* No version in .csproj

* added flag to disable training code (#227)

* Exit if perfect model produced (#220)

* removed header (#228)

* removed header

* added auto generated header

* removed console read key (#229)

* Fix model path in generated file (#230)

* removed console read key

* fix model path

* fix test

* reorder samples (#231)

* remove rule that infers column purpose as categorical if # of distinct values is < 100 (#233)

* Null reference exception fix for finding best model when some runs have failed (#239)

* samples fixes (#238)

* fix for defaulting Averaged Perceptron # of iterations to 10 (#237)

* Bug bash feedback Feb 27. API changes and sample changes (#240)

* Bug bash feedback Feb 27. 
API changes 
Sample changes
Exception fix

* Samples / API rev from 2/27 bug bash feedback (#242)

* changed the directory structure for generated project (#243)

* changed the directory structure for generated project

* changed test

* upgraded commandline package

* Fix test file locations on OSX (#235)

* fix test file locations on OSX

* changing to Path.Combine()

* Additional Path.Combine()

* Remove ConsoleCodeGeneratorTests.GeneratedTrainCodeTest.received.txt

* Additional Path.Combine()

* add back in double comparison fix

* remove metrics agent NaN returns

* test fix

* test format fix

* mock out path

Thanks to @daholste for additional fixes!

* upgrade to latest ML.NET public surface (#246)

* Upgrade to ML.NET 0.11 (#247)

* initial changes

* fix lightgbm

* changed normalize method

* added tests

* fix tests

* fix test

* Private preview final API changes (#250)

* .NET framework design guidelines applied to public surface
* WhitelistedTrainers -> Trainers

* Add estimator to public API iteration result (#248)

* LightGBM pipeline serialization fix (#251)

* Change order that we search for TextLoader's parameters (#256)

* CLI IFileInfo null exception fix (#254)

* Averaged Perceptron pipeline serialization fix (#257)

* Upgrade command-line-api and default folder name change (#258)

* change in defautl folderName

* upgrade command line

* Update src/mlnet/Program.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* eliminate IFileInfo from CLI (#260)

* Rev samples towards private preview; ignored columns fix (#259)

* remove unused methods in consolehelper and nit picks in generated code (#261)

* nit picks

* change in console helper

* fix tests

* add space

* fix tests

* added nuget sources in generated csproj (#262)

* added nuget sources in csproj

* changed the structure in generated code

* space

* upgrade to mlnet 0.11 (#263)

* Formatting CLI metrics (#264)

Ensures space between printed metrics (also model counter). Right aligned metrics. Extended AUC to four digits.

* Add implementation of non -ova multi class trainers code gen (#267)

* added non ova multi class learners

* added tests

* test cases

* Add caching (#249)

* AdvancedExperimentSettings sample nits (#265)

* Add sampling key column (#268)

* Initial work for multi-class classification support for CLI (#226)

* Initial work for multi-class classification support for CLI

* String updates

* more strings

* Whitelist non-OVA multi-class learners

* Refactor the orchestration of AutoML calls (#272)

* Do not auto-group columns with suggested purpose = 'Ignore' (#273)

* Fix: during type inferencing, parse whitespace strings as NaN (#271)

* Printing additional metrics in CLI for binary classification (#274)

* Printing additional metrics in CLI for binary classification

* Update src/mlnet/Utilities/ConsolePrinter.cs

* Add API option to store models on disk (instead of in memory); fix IEstimator memory leak (#269)

* Print failed iterations in CLI (#275)

* change the type to float from double (#277)

* cache arg implementation in CLI (#280)

* cache implementation

* corrected the null case

* added tests for all cases

* Remove duplicate value-to-key mapping transform for multiclass string labels (#283)

* Add post-trainer transform SDK infra; add KeyToValueMapping transform to CLI; fix: for generated multiclass models, convert predicted label from key to original label column type (#286)

* Implement ignore columns command line arg (#290)

* normalize line endings

* added --ignore-columns

* null checks

* unit tests

* Print winning iteration and runtime in CLI (#288)

* Print best metric and runtime

* Print best metric and runtime

* Line endings in AutoMLEngine.cs

* Rename time column to duration to match Python SDK

* Revert to MicroAccuracy and MacroAccuracy spellings

* Revert spelling of BinaryClassificationMetricsAgent to BinaryMetricsAgent to reduce merge conflicts

* Revert spelling of MulticlassMetricsAgent to MultiMetricsAgent to reduce merge conflicts

* missed some files

* Fix merge conflict

* Update AutoMLEngine.cs

* Add MacOS & Linux to CI; MacOS & Linux test fixes (#293)

* MicroAccuracy as default for multi-class (#295)

Change default optimization metric for multi-class classification to MicroAccuracy (accuracy). Previously it was set to MacroAccuracy.

* Null exception for ignorecolumns in CLI (#294)

* Null exception for ignorecolumns in CLI

* Check if ignore-columns array has values (as the default is now a empty array)

* Emit caching flag in pipeline object model. (Includes SuggestedPipelineBuilder refactor & debug string fixes / refactor) (#296)

* removed sln (#297)

* Caching enabling in code gen part -2 (#298)

* add

* added caching codegen

* support comma separated values for --ignore-columns (#300)

* default initialization for ignore columns (#302)

* default initialization

* adde null check

* Codegen for multiclass non-ova (#303)

* changes to template

* multicalss codegen

* test cases

* fix test cases

* Generated Project new structure. (#305)

* added new templates

* writing files to disck

* change path

* added new templates

* misisng braces

* fix bugs

* format code

* added util methods for solution file creation and addition of projects to it

* added extra packages to project files

* new tests

* added correct path for sln

* build fix

* fix build

* include using system in prediction class (#307)

* added using

* fix test

* Random number generator is not thread safe (#310)

* Random number generator is not thread safe

* Another local random generator

* Missed a few references

* Referncing AutoMlUtils.random instead of a local RNG

* More refs to mail RNG; remove Float as per https://github.com/dotnet/machinelearning/issues/1669

* Missed Random.cs

* Fix multiclass code gen (#314)

* compile error in codegen

* removes scores printing

* fix bugs

* fix test

* Fix compile error in codegen project (#319)

* removed redundant code

* fix test case

* Rev OVA pipeline node SDK output: wrap binary trainers as children inside parent OVA node (#317)

* Ova Multi class codegen support (#321)

* dummy

* multiova implementation

* fix tests

* remove inclusion list

* fix tests and console helper

* Rev run result trainer name for OVA: output different trainer name for each OVA + binary learner combination (#322)

* Rev run result trainer name for Ova: output different trainer name for each Ova + binary learner combination

* test fixes

* Console helper bug in generated code for multiclass (#323)

* fix

* fix test

* looping perlogclass

* fix test

* Initial version of Progress bar impl and CLI UI experience (#325)

* progressbar

* added progressbar and refactoring

* reverted

* revert sign assembly

* added headers and removed exception rethrow

* Setting model directory to temp directory (#327)

* Suggested changes to progress bar (#335)

* progressbar

* added progressbar and refactoring

* reverted

* revert sign assembly

* added headers and removed exception rethrow

* bug fixes and updates to UI

* added friendly name printing for metric

* formatting

* Rev Samples (#334)

* Telemetry2 (#333)

* Create test.txt

* Create test.txt

* changes needed for benchmarking

* forgot one file

* merge conflict fix

* fix build break

* back out my version of the fix for Label column issue and fix the original fix

* bogus file removal

* undo SuggestedPipeline change

* remove labelCol from pipeline suggester

* fix build break

* rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline)

* tweak queue in vsts-ci.yml

* CLI telemetry implementation

* Telemetry implementation

* delete unnecessary file and change file size bucket to actually log log2 instead of nearest ceil value

* add headers, remove comments

* one more header missing

* Fix progress bar in linux/osx (#336)

* progressbar

* added progressbar and refactoring

* reverted

* revert sign assembly

* added headers and removed exception rethrow

* bug fixes and updates to UI

* added friendly name printing for metric

* formatting

* change from task to thread

* Update src/mlnet/CodeGenerator/CodeGenerationHelper.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* Mem leak fix (#328)

* Create test.txt

* Create test.txt

* changes needed for benchmarking

* forgot one file

* merge conflict fix

* fix build break

* back out my version of the fix for Label column issue and fix the original fix

* bogus file removal

* undo SuggestedPipeline change

* remove labelCol from pipeline suggester

* fix build break

* rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline)

* tweak queue in vsts-ci.yml

* there is still investigation to be done but this fix works and solves memory leak problems

* minor refactor

* Upgrade ML.NET package (#343)

* Add cross-validation (CV), and auto-CV for small datasets; push common API experiment methods into base class (#287)

* restore old yml for internal pipeline so we can publish nuget again to devdiv stream (#344)

* Polishing the CLI UI part-1 (#338)

* formatting of pbar message

* Polishing the UI

* optimization

* rename variable

* Update src/mlnet/AutoML/AutoMLEngine.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* Update src/mlnet/CodeGenerator/CodeGenerationHelper.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* new message

* changed hhtp to https

* added iteration num + 1

* change string name and add color to artifacts

* change the message

* build errors

* added null checks

* added exception messsages to log file

* added exception messsages to log file

* CLI ML.NET version upgrade (#345)

* Sample revs; ColumnInformation property name revs; pre-featurizer fixes (#346)

* CLI -- consume logs from AutoML SDK (#349)

* Rename RunDetails --> RunDetail (#350)

* command line api upgrade and progress bar rendering bug (#366)

* added fix for all platforms progress bar

* upgrade nuget

* removed args from writeline

* change in the version (#368)

* fix few bugs in progressbar and verbosity (#374)

* fix few bugs in progressbar and verbosity

* removed unused name space

* Fix for folders with space in it while generating project (#376)

* support for folders with spaces

* added support for paths with space

* revert file

* change name of var

* remove spaces

* SMAC fix for minimizing metrics (#363)

* Formatting Regression metrics and progress bar display days. (#379)

* added progress bar day display and fix regression metrics

* fix formatting

* added total time

* formatted total time

* change command name and add pbar message (#380)

* change command name and add pbar message

* fix tests

* added aliases

* duplicate alias

* added another alias for task

* UI missing features (#382)

* added formatting changes

* added accuracy specifically

* downgrade the codepages (#384)

* Change in project structure (#385)

* initial changes

* Change in project structure

* correcting test

* change variable name

* fix tests

* fix tests

* fix more tests

* fix codegen errors

* adde log file message

* changed name of args

* change variable names

* fix test

* FileSizeBuckets in correct units (#387)

* Minor telemetry change to log in correct units and make our life easier in the future

* Use Ceiling instead of Round

* changed order (#388)

* prep work to transfer to ml.net (#389)

* move test projects to top level test subdir

* rename some projects to make naming consistent and make it build again

* fix test project refs

* Add AutoML components to build, fix issues related to that so it builds

* fix test cases, remove AppInsights ref from AutoML (#3329)

* [AutoML] disable netfx build leg for now (#3331)

* disable netfx build leg for now

* disable netfx build leg for now.

* [AutoML] Add AutoML XML documentation to all public members; migrate AutoML projects & tests into ML.NET solution; AutoML test fixes (#3351)

* [AutoML] Rev AutoML public API; add required native references to AutoML projects (#3364)

* [AutoML] Minor changes to generated project in CLI based on feedback (#3371)

* nitpicks for generated project

* revert back the target framework

* [AutoML] Migrate AutoML back to its own solution, w/ NuGet dependencies (#3373)

* Migrate AutoML back to its own solution, w/ NuGet dependencies

* build project updates; parameter name revert

* dummy change

* Revert "dummy change"

This reverts commit 3e8574266f556a4d5b6805eb55b4d8b8b84cf355.

* [AutoML] publish AutoML package (#3383)

* publish AutoML package

* Only leave automl and mlnet tests to run

* publish AutoML package

* Only leave automl and mlnet tests to run

* fix build issues when ml.net is not building

* bump version to 0.3 since that's the one we're going to ship for build (#3416)

* [AutoML] temporarily disable all but x64 platforms -- don't want to do native builds and can't find a way around that with the current VSTS pipeline (#3420)

* disable steps but keep phases to keep vsts build pipeline happy (#3423)

* API docs for experimentation (#3484)

* fixed path bug and regression metrics correction (#3504)

* changed the casing of option alias as it conflicts with --help (#3554)

* [AutoML] Generated project - FastTree nuget package inclusion dynamically (#3567)

* added support for fast tree nuget pack inclusion in generated project

* fix testcase

* changed the tool name in telemetry message

* dummy commit

* remove space

* dummy commit to trigger build

* [AutoML] Add AutoML example code (#3458)

* AutoML PipelineSuggester: don't recommend pipelines from first-stage trainers that failed (#3593)

* InferColumns API: Validate all columns specified in column info exist in inferred data view (#3599)

* [AutoML] AutoML SDK API: validate schema types of input IDataView (#3597)

* [AutoML] If first three iterations all fail, short-circuit AutoML experiment (#3591)

* mlnet CLI nupkg creation/signing (#3606)

* mlnet CLI nupkg creation/signing

* relmove includeinpackage from mlnet csproj

* address PR comments -- some minor reshuffling of stuff

* publish symbols for mlnet CLI

* fix case in NLog.config

* [AutoML] rename Auto to AutoML in namespace and nuget (#3609)

* mlnet CLI nupkg creation/signing

* [AutoML] take dependency on a specific ml.net version (#3610)

* take dependency on a specific ml.net version

* catch up to spelling fix for OptimizationTolerance

* force a specific ml.net nuget version, fix typo (#3616)

* [AutoML] Fix error handling in CLI.  (#3618)

* fix error handling

* renaming variables

* [AutoML] turn off line pragmas in .tt files to play nice with signing (#3617)

* turn off line pragmas in .tt files to play nice with signing

* dedupe tags

* change the param name (#3619)

* [AutoML]  return null instead of null ref crash on Model property accessor (#3620)

* return null instead of null ref crash on Model property accessor

* [AutoML] Handling label column names which have space and exception logging (#3624)

* fix case of label with space and exception logging

* final handler

* revert file

* use Name instead of FullName for telemetry filename hash (#3633)

* renamed classes (#3634)

* change ML.NET dependency to 1.0 (#3639)

[AutoML] undo pinning ML.NET dependency

* set exploration time default in CLI to half hour (#3640)

* [AutoML] step 2 of removing pinned nupkg versions (#3642)

* InferColumns API that consumes label column index -- Only rename label column to 'Label' for headerless files (#3643)

* [AutoML] Upgrade ml.net package in generated code (#3644)

* upgrade the mlnet package in gen code

* Update src/mlnet/Templates/Console/ModelProject.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* Update src/mlnet/Templates/Console/ModelProject.tt

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* added spaces

* [AutoML] Early stopping in CLI based on the exploration time (#3641)

* early stopping in CLI

* remove unused variables

* change back to thread

* remove sleep

* fix review comments

* remove ununsed usings

* format message

* collapse declaration

* remove unused param

* added environment.exit and removal of error message

* correction in message

* secs-> seconds

* exit code

* change value to 1

* reverse the declaration

* [AutoML] Change wording for CouldNotFinshOnTime message (#3655)

* set exploration time default in CLI to half hour

* [AutoML] Change wording for CouldNotFinshOnTime message

* [AutoML] Change wording for CouldNotFinshOnTime message

* even better wording for CouldNotFinshOnTime

* temp change to get around vsts publish failure (#3656)

* [AutoML] bump version to 0.4.0 (#3658)

* implement culture invariant strings (#3725)

* reset culture (#3730)

* [AutoML] Cross validation fixes; validate empty training / validation input data (#3794)

* [AutoML] Enable style cop rules & resolve errors (#3823)

* add task agnostic wrappers for autofit calls (#3860)

* [AutoML] CLI telemetry rev (#3789)

* delete automl .sln

* CLI -- regenerate templated CS files (#3954)

* [AutoML] Bump ML.NET package version to 1.2.0 in AutoML API and CLI; and AutoML package versions to 0.14.0 (#3958)

* Build AutoML NuGet package (#3961)

* Increment AutoML build version to 0.15.0 for preview. (#3968)

* added culture independent parsing (#3731)

* - convert tests to xunit
- take project level dependency on ML.NET components instead of nuget
- set up bestfriends relationship to ML.Core and remove some of the copies of util classes from AutoML.NET (more work needed to fully remove them, work item 4064)
- misc build script changes to address PR comments

* address issues only showing up in a couple configurations during CI build

* fix cut&paste error

* [AutoML] Bump version to ML.NET 1.3.1 in AutoML API and CLI and AutoML package version to 0.15.1 (#4071)

* bumped version

* change versions in nupkg

* revert version bump in branch props

* [AutoML] Fix for Exception thrown in cross val when one of the score equals infinity. (#4073)

* bumped version

* change versions in nupkg

* revert version bump in branch props

* added infinity fix

* changes signing (#4079)

* Addressed PR comments and build issues
- sync block on creating test data file (failed intermittently)
- removed classes we copied over from ML.Core and fixed their uses to de-dupe and use original ML.Core versions since we now have InternalsVisible and BestFriends
- Fixed nupkg creation  to use projects insted of public nuget version for AutoML
- Fixed a bunch of unit tests that didn't actually test what they were supposed to test, while removing cut&past code and dependencies.
- Few more misc small changes

* minor nit - removed unused folder ref

* Fix the .sln file for the right configurations.

* Fix mistake in .sln file

* test fixes and disable one test

* fix tests, re-add AutoML samples csproj

* bumped VS version to 16 in .sln, removed InternalsVisible for a dead assembly, removed unused references from AutoML test project

* Updated docs to include PredictedLabel member (#4107)

* Fixed build errors resulting from upgrade to VS2019 compilers

* Added additional message describing the previous fix

* Updated docs to include PredictedLabel member

* Added CODEOWNERS file in the .github/ folder. (#4140)

* Added CODEOWNERS file in the .github/ folder. This allows reviewers to review any changes in the machine learning repository

* Updated .github/CODEOWNERS with the team instead of individual reviewers

* Added AutoML team reviewers (#4144)

* Added CODEOWNERS file in the .github/ folder. This allows reviewers to review any changes in the machine learning repository

* Updated .github/CODEOWNERS with the team instead of individual reviewers

* Added AutoML team reviwers to files owned by AutoML team

* Added AutoML team reviwers to files owned by AutoML team

* Removed two files that don't exist for AutoML team in CODEOWNERS

* Build extension method to reload changes without specifying model name (#4146)

* Image classification preview 2. (#4151)

* Image classification preview 2.

* PR feedback.

* Add unit-test.

* Add unit-test.

* Add unit-test.

* Add unit-test.

* Use Path.Combine instead of Join.

* fix test dataset path.

* fix test dataset path.

* Improve test.

* Improve test.

* Increase epochs in tests.

* Disable test on Ubuntu.

* Move test to its own project.

* Move test to its own project.

* Move test to its own project.

* Move test to its own file.

* cleanup.

* Disable parallel execution of tensorflow tests.

* PR feedback.

* PR feedback.

* PR feedback.

* PR feedback.

* Prevent TF test to execute in parallel.

* PR feedback.

* Build error.

* clean up.

* Added export functionality for LpNormNormalizingTransformer

* Syncing upstream fork (#11)

* Throw error on incorrect Label name in InferColumns API (#47)

* Added sequential grouping of columns

* reverted the file

* addded infer columns label name checking

* added column detection error

* removed unsed usings

* added quotes

* replace Where with Any clause

* replace Where with Any clause

* Set Nullable Auto params to null values (#50)

* Added sequential grouping of columns

* reverted the file

* added auto params as null

* change to the update fields method

* First public api propsal (#52)

* Includes following
1) Final proposal for 0.1 public API surface
2) Prefeaturization
3) Splitting train data into train and validate when validation data is null
4) Providing end to end samples one each for regression, binaryclassification and multiclass classification

* Incorporating code review feedbacks

* Revert "Set Nullable Auto params to null values" (#53)

* Revert "First public api propsal (#52)"

This reverts commit e4a64cf4aeab13ee9e5bf0efe242da3270241bd7.

* Revert "Set Nullable Auto params to null values (#50)"

This reverts commit 41c663cd14247d44022f40cf2dce5977dbab282d.

* AutoFit return type is now an IEnumerable (#55)

AutoFit returns is now an IEnumerable - this enables many good things

Implementing variety of early stopping criteria (See sample)
Early discard of models that are no good. This improves memory usage efficiency. (See sample)
No need to implement a callback to get results back
Getting best score is now outside of API implementation. It is a simple math function to compare scores (See sample).

Also templatized the return type for better type safety through out the code.

* misc fixes & test additions, towards 0.1 release (#56)

* Enable UnitTests on build server (#57)

* 1) Making trainer name public (#62)

2) Fixing up samples to reflect it

*  Initial version of CLI tool for mlnet (#61)

* added global tool initial project

* removed unneccesary files, renamed files

* refactoring and added base abstract classes for trainer generator

* removed unused class

* Added classes for transforms

* added transform generate dummy classes

* more refactoring, added first transform

* more refactoring and added classes

* changed the project structure

* restructing added options class

* sln changes

* refactored options to different class:

* added more logic for code generation of class

* misc changes

* reverted file

* added commandline api package

* reverted sample

* added new command line api parser

* added normalization of column names

* Added command defaults and error message

* implementation of all trainers

* changed auto to null

* added all transform generators

* added error handling when args is empty and minor changes due to change in AutoML api names

* changed the name of param

* added new command line options and restructuring code

* renamed proj file and added solution

* Added code to generate usings, Fixed few bugs in the code

* added validation to the command line options

* changed project name

* Bug fixes due to API change in AutoML

* changed directory structure

* added test framework and basic tests

* added more tests

* added improvements to template and error handling

* renamed the estimator name

* fixed test case

* added comments

* added headers

* changed namespace and removed unneccesary properties from project

* Revert "changed namespace and removed unneccesary properties from project"

This reverts commit 9edae033e9845e910f663f296e168f1182b84f5f.

* fixed test cases and renamed namespaces

* cleaned up proj file

* added folder structure

* added symbols/tokens for strings

* added more tests

* review comments

* modified test cases

* review comments

* change in the exception message

* normalized line endings

* made method private static

* simplified range building /optimization

* minor fix

* added header

* added static methods in command where necessary

* nit picks

*  made few methods static

* review comments

* nitpick

* remove line pragmas

* fix test case

* Use better AutiFit overload and ignore Multiclass (#64)

* Upgrading CLI to produce ML.NET V.10 APIs and bunch of Refactoring tasks (#65)

* Added sequential grouping of columns

* reverted the file

* upgrade to v .10 and refactoring

* added null check

* fixed unit tests

* review comments

* removed the settings change

* added regions

* fixed unit tests

* Upgrade ML.NET package to 0.10.0 (#70)

* Change in template to accomodate new API of TextLoader (#72)

* Added sequential grouping of columns

* reverted the file

* changed to new API of Text Loader

* changed signature

* added params for taking additional settings

* changes to codegen params

* refactoring of templates and fixing errors

* Enable gated check for mlnet.tests (#79)

* Added sequential grouping of columns

* reverted the file

* changed to new API of Text Loader

* changed signature

* added params for taking additional settings

* changes to codegen params

* refactoring of templates and fixing errors

* added run-tests.proj and referred it in build.proj

* CLI tool - make validation dataset optional and support for crossvalidation in generated code (#83)

* Added sequential grouping of columns

* reverted the file

* bug fixes, more logic to templates to support cross-validate

* formatting and fix type in consolehelper

* Added logic in templates

* revert settings

* benchmarking related changes (#63)

* Create test.txt

* Create test.txt

* changes needed for benchmarking

* forgot one file

* merge conflict fix

* fix build break

* back out my version of the fix for Label column issue and fix the original fix

* bogus file removal

* undo SuggestedPipeline change

* remove labelCol from pipeline suggester

* fix build break

* fix fast forest learner (don't sweep over learning rate) (#88)

* Made changes to Have non-calibrated scoring for binary classifiers (#86)

* Added sequential grouping of columns

* reverted the file

* added calibration workaround

* removed print probability

* reverted settings

* rev ColumnInference API: can take label index; rev output object types; add tests (#89)

* rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (#99)

* Create test.txt

* Create test.txt

* changes needed for benchmarking

* forgot one file

* merge conflict fix

* fix build break

* back out my version of the fix for Label column issue and fix the original fix

* bogus file removal

* undo SuggestedPipeline change

* remove labelCol from pipeline suggester

* fix build break

* rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline)

* publish nuget (#101)

* use dotnet-internal-temp agent for internal build

* use dotnet-internal feed

* Fix Codegen for columnConvert and ValueToKeyMapping transform and add individual transform tests (#95)

* Added sequential grouping of columns

* reverted the file

* fix usings for type convert

* added transforms tests

* review comments

* When generating usings choose only distinct usings directives (#94)

* Added sequential grouping of columns

* reverted the file

* Added code to have unique strings

* refactoring

* minor fix

* minor fix

* Autofit overloads + cancellation + progress callbacks

1) Introduce AutoFit overloads (basic and advanced)
2) AutoFit Cancellation
3) AutoFit progress callbacks

* Default the kfolds to value 5 in CLI generated code (#115)

* Added sequential grouping of columns

* reverted the file

* Set up CI with Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* remove file

* added kfold param and defaulted to value

* changed type

* added for regression

* Remove extra ; from generated code (#114)

* Added sequential grouping of columns

* reverted the file

* Set up CI with Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* removed extra ; from generated code

* removed file

* fix unit tests

* TimeoutInSeconds (#116)

Specifying timeout in seconds instead of minutes

* Added more command line args implementation to CLI tool and refactoring (#110)

* Added sequential grouping of columns

* reverted the file

* Set up CI with Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* added git status

* reverted change

* added codegen options and refactoring

* minor fixes'

* renamed params, minor refactoring

* added tests for commandline and refactoring

* removed file

* added back the test case

* minor fixes

* Update src/mlnet.Test/CommandLineTests.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* review comments

*  capitalize the first character

* changed the name of test case

* remove unused directives

* Fail gracefully if unable to instantiate data view with swept parameters (#125)

* gracefully fail if fail to parse a datai

* rev

* validate AutoFit 'Features' column must be of type R4 (#132)

* Samples: exceptions / nits (#124)

* Logging support in CLI + Implementation of cmd args [--name,--output,--verbosity] (#121)

* addded logging and helper methods

* fixing code after merge

* added resx files, added logger framework, added logging messages

* added new options

* added spacing

* minor fixes

* change command description

* rename option, add headers, include new param in test

* formatted

* build fix

*  changed option name

* Added NlogConfig file

* added back config package

* fix tests

* added correct validation check (#137)

* Use CreateTextLoader<T>(..)  instead of CreateTextLoader(..) (#138)

* added support to loaddata by class in the generated code

* fix tests

* changed CreateTextLoader to ReadFromTextFile method. (#140)

* changed textloader to readfromtextfile method

* formatting

* exception fixes (#136)

* infer purpose of hidden columns as 'ignore' (#142)

* Added approval tests and bunch of refactoring of code and normalizing namespaces (#148)

* changed textloader to readfromtextfile method

* formatting

* added approval tests and refactoring of code

* removed few comments

* API 2.0 skeleton (#149)

Incorporating API review feedback

* The CV code should come before the training when there is no test dataset in generated code (#151)

* reorder cv code

* build fix

* fixed structure

* Format the generated code + bunch of misc tasks (#152)

* added formatting and minor changes for reordering cv

* fixing the template

* minor changes

* formatting changes

* fixed approval test

* removed unused nuget

* added missing value replacing

* added test for new transform

* fix test

* Update src/mlnet/Templates/Console/MLCodeGen.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* Sanitize the column names in CLI (#162)

* added sanitization layer in CLI

* fix test

* changed exception.StackTrace to exception.ToString()

* fix package name (#168)

* Rev public API (#163)

* Rename TransformGeneratorBase .cs to TransformGeneratorBase.cs (#153)

* Fix minor version for the repository + remove Nlog config package (#171)

*  changed the minor version

* removed the nlog config package

* Added new test to columninfo and fixing up API (#178)

* Make optimizing metric customizable and add trainer whitelist functionality (#172)

* API rev (#181)

* propagate root MLContext thru AutoML (instead of creating our own) (#182)

* Enabling new command line args (#183)

* fix package name

* initial commit

* added more commandline args

* fixed tests

* added headers

* fix tests

* fix test

* rename 'AutoFitter' to 'Experiment' (#169)

* added tests (#187)

* rev InferColumns to accept ColumnInfo input param (#186)

* Implement argument --has-header and change usage of dataset (#194)

* added has header and fixed dataset and train dataset

* fix tests

* removed dummy command (#195)

* Fix bug for regression and sanitize input label from user (#198)

* removed dummy command

* sanitize label and fix template

* fix tests

* Do not generate code concatenating columns when the dataset has a single feature column (#191)

* Include some missed logging in the generated code.  (#199)

* added logging messages for generated code

* added log messages

* deleted file

* cleaning up proj files (#185)

* removed platform target

* removed platform target

* Some spaces and extra lines + bug in output path  (#204)

* nit picks

* nit picks

* fix test

* accept label from user input and provide in generated code (#205)

* Rev handling of weight / label columns (#203)

* migrate to private ML.NET nuget for latest bug fixes (#131)

* fix multiclass with nonstandard label (#207)

* Multiclass nondefault label test (#208)

* printing escaped chars + bug (#212)

* delete unused internal samples (#211)

* fix SMAC bug that causes multiclass sample to infinite loop (#209)

* Rev user input validation for new API (#210)

* added console message for exit and nit picks (#215)

* exit when exception encountered (#216)

* Seal API classes (and make EnableCaching internal) (#217)

* Suggested sample nits (feel free to ask for any of these to be reverted) (#219)

* User input column type validation (#218)

* upgrade commandline and renaming (#221)

* upgrade commandline and renaming

* renaming fields

* Make build.sh, init-tools.sh, & run.sh executable on OSX/Linux (#225)

*  CLI argument descriptions updated (#224)

* CLI argument descriptions updated

* No version in .csproj

* added flag to disable training code (#227)

* Exit if perfect model produced (#220)

* removed header (#228)

* removed header

* added auto generated header

* removed console read key (#229)

* Fix model path in generated file (#230)

* removed console read key

* fix model path

* fix test

* reorder samples (#231)

* remove rule that infers column purpose as categorical if # of distinct values is < 100 (#233)

* Null reference exception fix for finding best model when some runs have failed (#239)

* samples fixes (#238)

* fix for defaulting Averaged Perceptron # of iterations to 10 (#237)

* Bug bash feedback Feb 27. API changes and sample changes (#240)

* Bug bash feedback Feb 27. 
API changes 
Sample changes
Exception fix

* Samples / API rev from 2/27 bug bash feedback (#242)

* changed the directory structure for generated project (#243)

* changed the directory structure for generated project

* changed test

* upgraded commandline package

* Fix test file locations on OSX (#235)

* fix test file locations on OSX

* changing to Path.Combine()

* Additional Path.Combine()

* Remove ConsoleCodeGeneratorTests.GeneratedTrainCodeTest.received.txt

* Additional Path.Combine()

* add back in double comparison fix

* remove metrics agent NaN returns

* test fix

* test format fix

* mock out path

Thanks to @daholste for additional fixes!

* upgrade to latest ML.NET public surface (#246)

* Upgrade to ML.NET 0.11 (#247)

* initial changes

* fix lightgbm

* changed normalize method

* added tests

* fix tests

* fix test

* Private preview final API changes (#250)

* .NET framework design guidelines applied to public surface
* WhitelistedTrainers -> Trainers

* Add estimator to public API iteration result (#248)

* LightGBM pipeline serialization fix (#251)

* Change order that we search for TextLoader's parameters (#256)

* CLI IFileInfo null exception fix (#254)

* Averaged Perceptron pipeline serialization fix (#257)

* Upgrade command-line-api and default folder name change (#258)

* change in defautl folderName

* upgrade command line

* Update src/mlnet/Program.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* eliminate IFileInfo from CLI (#260)

* Rev samples towards private preview; ignored columns fix (#259)

* remove unused methods in consolehelper and nit picks in generated code (#261)

* nit picks

* change in console helper

* fix tests

* add space

* fix tests

* added nuget sources in generated csproj (#262)

* added nuget sources in csproj

* changed the structure in generated code

* space

* upgrade to mlnet 0.11 (#263)

* Formatting CLI metrics (#264)

Ensures space between printed metrics (also model counter). Right aligned metrics. Extended AUC to four digits.

* Add implementation of non -ova multi class trainers code gen (#267)

* added non ova multi class learners

* added tests

* test cases

* Add caching (#249)

* AdvancedExperimentSettings sample nits (#265)

* Add sampling key column (#268)

* Initial work for multi-class classification support for CLI (#226)

* Initial work for multi-class classification support for CLI

* String updates

* more strings

* Whitelist non-OVA multi-class learners

* Refactor the orchestration of AutoML calls (#272)

* Do not auto-group columns with suggested purpose = 'Ignore' (#273)

* Fix: during type inferencing, parse whitespace strings as NaN (#271)

* Printing additional metrics in CLI for binary classification (#274)

* Printing additional metrics in CLI for binary classification

* Update src/mlnet/Utilities/ConsolePrinter.cs

* Add API option to store models on disk (instead of in memory); fix IEstimator memory leak (#269)

* Print failed iterations in CLI (#275)

* change the type to float from double (#277)

* cache arg implementation in CLI (#280)

* cache implementation

* corrected the null case

* added tests for all cases

* Remove duplicate value-to-key mapping transform for multiclass string labels (#283)

* Add post-trainer transform SDK infra; add KeyToValueMapping transform to CLI; fix: for generated multiclass models, convert predicted label from key to original label column type (#286)

* Implement ignore columns command line arg (#290)

* normalize line endings

* added --ignore-columns

* null checks

* unit tests

* Print winning iteration and runtime in CLI (#288)

* Print best metric and runtime

* Print best metric and runtime

* Line endings in AutoMLEngine.cs

* Rename time column to duration to match Python SDK

* Revert to MicroAccuracy and MacroAccuracy spellings

* Revert spelling of BinaryClassificationMetricsAgent to BinaryMetricsAgent to reduce merge conflicts

* Revert spelling of MulticlassMetricsAgent to MultiMetricsAgent to reduce merge conflicts

* missed some files

* Fix merge conflict

* Update AutoMLEngine.cs

* Add MacOS & Linux to CI; MacOS & Linux test fixes (#293)

* MicroAccuracy as default for multi-class (#295)

Change default optimization metric for multi-class classification to MicroAccuracy (accuracy). Previously it was set to MacroAccuracy.

* Null exception for ignorecolumns in CLI (#294)

* Null exception for ignorecolumns in CLI

* Check if ignore-columns array has values (as the default is now a empty array)

* Emit caching flag in pipeline object model. (Includes SuggestedPipelineBuilder refactor & debug string fixes / refactor) (#296)

* removed sln (#297)

* Caching enabling in code gen part -2 (#298)

* add

* added caching codegen

* support comma separated values for --ignore-columns (#300)

* default initialization for ignore columns (#302)

* default initialization

* adde null check

* Codegen for multiclass non-ova (#303)

* changes to template

* multicalss codegen

* test cases

* fix test cases

* Generated Project new structure. (#305)

* added new templates

* writing files to disck

* change path

* added new templates

* misisng braces

* fix bugs

* format code

* added util methods for solution file creation and addition of projects to it

* added extra packages to project files

* new tests

* added correct path for sln

* build fix

* fix build

* include using system in prediction class (#307)

* added using

* fix test

* Random number generator is not thread safe (#310)

* Random number generator is not thread safe

* Another local random generator

* Missed a few references

* Referncing AutoMlUtils.random instead of a local RNG

* More refs to mail RNG; remove Float as per https://github.com/dotnet/machinelearning/issues/1669

* Missed Random.cs

* Fix multiclass code gen (#314)

* compile error in codegen

* removes scores printing

* fix bugs

* fix test

* Fix compile error in codegen project (#319)

* removed redundant code

* fix test case

* Rev OVA pipeline node SDK output: wrap binary trainers as children inside parent OVA node (#317)

* Ova Multi class codegen support (#321)

* dummy

* multiova implementation

* fix tests

* remove inclusion list

* fix tests and console helper

* Rev run result trainer name for OVA: output different trainer name for each OVA + binary learner combination (#322)

* Rev run result trainer name for Ova: output different trainer name for each Ova + binary learner combination

* test fixes

* Console helper bug in generated code for multiclass (#323)

* fix

* fix test

* looping perlogclass

* fix test

* Initial version of Progress bar impl and CLI UI experience (#325)

* progressbar

* added progressbar and refactoring

* reverted

* revert sign assembly

* added headers and removed exception rethrow

* Setting model directory to temp directory (#327)

* Suggested changes to progress bar (#335)

* progressbar

* added progressbar and refactoring

* reverted

* revert sign assembly

* added headers and removed exception rethrow

* bug fixes and updates to UI

* added friendly name printing for metric

* formatting

* Rev Samples (#334)

* Telemetry2 (#333)

* Create test.txt

* Create test.txt

* changes needed for benchmarking

* forgot one file

* merge conflict fix

* fix build break

* back out my version of the fix for Label column issue and fix the original fix

* bogus file removal

* undo SuggestedPipeline change

* remove labelCol from pipeline suggester

* fix build break

* rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline)

* tweak queue in vsts-ci.yml

* CLI telemetry implementation

* Telemetry implementation

* delete unnecessary file and change file size bucket to actually log log2 instead of nearest ceil value

* add headers, remove comments

* one more header missing

* Fix progress bar in linux/osx (#336)

* progressbar

* added progressbar and refactoring

* reverted

* revert sign assembly

* added headers and removed exception rethrow

* bug fixes and updates to UI

* added friendly name printing for metric

* formatting

* change from task to thread

* Update src/mlnet/CodeGenerator/CodeGenerationHelper.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* Mem leak fix (#328)

* Create test.txt

* Create test.txt

* changes needed for benchmarking

* forgot one file

* merge conflict fix

* fix build break

* back out my version of the fix for Label column issue and fix the original fix

* bogus file removal

* undo SuggestedPipeline change

* remove labelCol from pipeline suggester

* fix build break

* rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline)

* tweak queue in vsts-ci.yml

* there is still investigation to be done but this fix works and solves memory leak problems

* minor refactor

* Upgrade ML.NET package (#343)

* Add cross-validation (CV), and auto-CV for small datasets; push common API experiment methods into base class (#287)

* restore old yml for internal pipeline so we can publish nuget again to devdiv stream (#344)

* Polishing the CLI UI part-1 (#338)

* formatting of pbar message

* Polishing the UI

* optimization

* rename variable

* Update src/mlnet/AutoML/AutoMLEngine.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* Update src/mlnet/CodeGenerator/CodeGenerationHelper.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* new message

* changed hhtp to https

* added iteration num + 1

* change string name and add color to artifacts

* change the message

* build errors

* added null checks

* added exception messsages to log file

* added exception messsages to log file

* CLI ML.NET version upgrade (#345)

* Sample revs; ColumnInformation property name revs; pre-featurizer fixes (#346)

* CLI -- consume logs from AutoML SDK (#349)

* Rename RunDetails --> RunDetail (#350)

* command line api upgrade and progress bar rendering bug (#366)

* added fix for all platforms progress bar

* upgrade nuget

* removed args from writeline

* change in the version (#368)

* fix few bugs in progressbar and verbosity (#374)

* fix few bugs in progressbar and verbosity

* removed unused name space

* Fix for folders with space in it while generating project (#376)

* support for folders with spaces

* added support for paths with space

* revert file

* change name of var

* remove spaces

* SMAC fix for minimizing metrics (#363)

* Formatting Regression metrics and progress bar display days. (#379)

* added progress bar day display and fix regression metrics

* fix formatting

* added total time

* formatted total time

* change command name and add pbar message (#380)

* change command name and add pbar message

* fix tests

* added aliases

* duplicate alias

* added another alias for task

* UI missing features (#382)

* added formatting changes

* added accuracy specifically

* downgrade the codepages (#384)

* Change in project structure (#385)

* initial changes

* Change in project structure

* correcting test

* change variable name

* fix tests

* fix tests

* fix more tests

* fix codegen errors

* adde log file message

* changed name of args

* change variable names

* fix test

* FileSizeBuckets in correct units (#387)

* Minor telemetry change to log in correct units and make our life easier in the future

* Use Ceiling instead of Round

* changed order (#388)

* prep work to transfer to ml.net (#389)

* move test projects to top level test subdir

* rename some projects to make naming consistent and make it build again

* fix test project refs

* Add AutoML components to build, fix issues related to that so it builds

* fix test cases, remove AppInsights ref from AutoML (#3329)

* [AutoML] disable netfx build leg for now (#3331)

* disable netfx build leg for now

* disable netfx build leg for now.

* [AutoML] Add AutoML XML documentation to all public members; migrate AutoML projects & tests into ML.NET solution; AutoML test fixes (#3351)

* [AutoML] Rev AutoML public API; add required native references to AutoML projects (#3364)

* [AutoML] Minor changes to generated project in CLI based on feedback (#3371)

* nitpicks for generated project

* revert back the target framework

* [AutoML] Migrate AutoML back to its own s…
harishsk added a commit that referenced this pull request Sep 11, 2019
* Fixed build errors resulting from upgrade to VS2019 compilers

* Added additional message describing the previous fix

* Syncing upstream fork (#10)

* Throw error on incorrect Label name in InferColumns API (#47)

* Added sequential grouping of columns

* reverted the file

* addded infer columns label name checking

* added column detection error

* removed unsed usings

* added quotes

* replace Where with Any clause

* replace Where with Any clause

* Set Nullable Auto params to null values (#50)

* Added sequential grouping of columns

* reverted the file

* added auto params as null

* change to the update fields method

* First public api propsal (#52)

* Includes following
1) Final proposal for 0.1 public API surface
2) Prefeaturization
3) Splitting train data into train and validate when validation data is null
4) Providing end to end samples one each for regression, binaryclassification and multiclass classification

* Incorporating code review feedbacks

* Revert "Set Nullable Auto params to null values" (#53)

* Revert "First public api propsal (#52)"

This reverts commit e4a64cf4aeab13ee9e5bf0efe242da3270241bd7.

* Revert "Set Nullable Auto params to null values (#50)"

This reverts commit 41c663cd14247d44022f40cf2dce5977dbab282d.

* AutoFit return type is now an IEnumerable (#55)

AutoFit returns is now an IEnumerable - this enables many good things

Implementing variety of early stopping criteria (See sample)
Early discard of models that are no good. This improves memory usage efficiency. (See sample)
No need to implement a callback to get results back
Getting best score is now outside of API implementation. It is a simple math function to compare scores (See sample).

Also templatized the return type for better type safety through out the code.

* misc fixes & test additions, towards 0.1 release (#56)

* Enable UnitTests on build server (#57)

* 1) Making trainer name public (#62)

2) Fixing up samples to reflect it

*  Initial version of CLI tool for mlnet (#61)

* added global tool initial project

* removed unneccesary files, renamed files

* refactoring and added base abstract classes for trainer generator

* removed unused class

* Added classes for transforms

* added transform generate dummy classes

* more refactoring, added first transform

* more refactoring and added classes

* changed the project structure

* restructing added options class

* sln changes

* refactored options to different class:

* added more logic for code generation of class

* misc changes

* reverted file

* added commandline api package

* reverted sample

* added new command line api parser

* added normalization of column names

* Added command defaults and error message

* implementation of all trainers

* changed auto to null

* added all transform generators

* added error handling when args is empty and minor changes due to change in AutoML api names

* changed the name of param

* added new command line options and restructuring code

* renamed proj file and added solution

* Added code to generate usings, Fixed few bugs in the code

* added validation to the command line options

* changed project name

* Bug fixes due to API change in AutoML

* changed directory structure

* added test framework and basic tests

* added more tests

* added improvements to template and error handling

* renamed the estimator name

* fixed test case

* added comments

* added headers

* changed namespace and removed unneccesary properties from project

* Revert "changed namespace and removed unneccesary properties from project"

This reverts commit 9edae033e9845e910f663f296e168f1182b84f5f.

* fixed test cases and renamed namespaces

* cleaned up proj file

* added folder structure

* added symbols/tokens for strings

* added more tests

* review comments

* modified test cases

* review comments

* change in the exception message

* normalized line endings

* made method private static

* simplified range building /optimization

* minor fix

* added header

* added static methods in command where necessary

* nit picks

*  made few methods static

* review comments

* nitpick

* remove line pragmas

* fix test case

* Use better AutiFit overload and ignore Multiclass (#64)

* Upgrading CLI to produce ML.NET V.10 APIs and bunch of Refactoring tasks (#65)

* Added sequential grouping of columns

* reverted the file

* upgrade to v .10 and refactoring

* added null check

* fixed unit tests

* review comments

* removed the settings change

* added regions

* fixed unit tests

* Upgrade ML.NET package to 0.10.0 (#70)

* Change in template to accomodate new API of TextLoader (#72)

* Added sequential grouping of columns

* reverted the file

* changed to new API of Text Loader

* changed signature

* added params for taking additional settings

* changes to codegen params

* refactoring of templates and fixing errors

* Enable gated check for mlnet.tests (#79)

* Added sequential grouping of columns

* reverted the file

* changed to new API of Text Loader

* changed signature

* added params for taking additional settings

* changes to codegen params

* refactoring of templates and fixing errors

* added run-tests.proj and referred it in build.proj

* CLI tool - make validation dataset optional and support for crossvalidation in generated code (#83)

* Added sequential grouping of columns

* reverted the file

* bug fixes, more logic to templates to support cross-validate

* formatting and fix type in consolehelper

* Added logic in templates

* revert settings

* benchmarking related changes (#63)

* Create test.txt

* Create test.txt

* changes needed for benchmarking

* forgot one file

* merge conflict fix

* fix build break

* back out my version of the fix for Label column issue and fix the original fix

* bogus file removal

* undo SuggestedPipeline change

* remove labelCol from pipeline suggester

* fix build break

* fix fast forest learner (don't sweep over learning rate) (#88)

* Made changes to Have non-calibrated scoring for binary classifiers (#86)

* Added sequential grouping of columns

* reverted the file

* added calibration workaround

* removed print probability

* reverted settings

* rev ColumnInference API: can take label index; rev output object types; add tests (#89)

* rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (#99)

* Create test.txt

* Create test.txt

* changes needed for benchmarking

* forgot one file

* merge conflict fix

* fix build break

* back out my version of the fix for Label column issue and fix the original fix

* bogus file removal

* undo SuggestedPipeline change

* remove labelCol from pipeline suggester

* fix build break

* rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline)

* publish nuget (#101)

* use dotnet-internal-temp agent for internal build

* use dotnet-internal feed

* Fix Codegen for columnConvert and ValueToKeyMapping transform and add individual transform tests (#95)

* Added sequential grouping of columns

* reverted the file

* fix usings for type convert

* added transforms tests

* review comments

* When generating usings choose only distinct usings directives (#94)

* Added sequential grouping of columns

* reverted the file

* Added code to have unique strings

* refactoring

* minor fix

* minor fix

* Autofit overloads + cancellation + progress callbacks

1) Introduce AutoFit overloads (basic and advanced)
2) AutoFit Cancellation
3) AutoFit progress callbacks

* Default the kfolds to value 5 in CLI generated code (#115)

* Added sequential grouping of columns

* reverted the file

* Set up CI with Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* remove file

* added kfold param and defaulted to value

* changed type

* added for regression

* Remove extra ; from generated code (#114)

* Added sequential grouping of columns

* reverted the file

* Set up CI with Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* removed extra ; from generated code

* removed file

* fix unit tests

* TimeoutInSeconds (#116)

Specifying timeout in seconds instead of minutes

* Added more command line args implementation to CLI tool and refactoring (#110)

* Added sequential grouping of columns

* reverted the file

* Set up CI with Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* added git status

* reverted change

* added codegen options and refactoring

* minor fixes'

* renamed params, minor refactoring

* added tests for commandline and refactoring

* removed file

* added back the test case

* minor fixes

* Update src/mlnet.Test/CommandLineTests.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* review comments

*  capitalize the first character

* changed the name of test case

* remove unused directives

* Fail gracefully if unable to instantiate data view with swept parameters (#125)

* gracefully fail if fail to parse a datai

* rev

* validate AutoFit 'Features' column must be of type R4 (#132)

* Samples: exceptions / nits (#124)

* Logging support in CLI + Implementation of cmd args [--name,--output,--verbosity] (#121)

* addded logging and helper methods

* fixing code after merge

* added resx files, added logger framework, added logging messages

* added new options

* added spacing

* minor fixes

* change command description

* rename option, add headers, include new param in test

* formatted

* build fix

*  changed option name

* Added NlogConfig file

* added back config package

* fix tests

* added correct validation check (#137)

* Use CreateTextLoader<T>(..)  instead of CreateTextLoader(..) (#138)

* added support to loaddata by class in the generated code

* fix tests

* changed CreateTextLoader to ReadFromTextFile method. (#140)

* changed textloader to readfromtextfile method

* formatting

* exception fixes (#136)

* infer purpose of hidden columns as 'ignore' (#142)

* Added approval tests and bunch of refactoring of code and normalizing namespaces (#148)

* changed textloader to readfromtextfile method

* formatting

* added approval tests and refactoring of code

* removed few comments

* API 2.0 skeleton (#149)

Incorporating API review feedback

* The CV code should come before the training when there is no test dataset in generated code (#151)

* reorder cv code

* build fix

* fixed structure

* Format the generated code + bunch of misc tasks (#152)

* added formatting and minor changes for reordering cv

* fixing the template

* minor changes

* formatting changes

* fixed approval test

* removed unused nuget

* added missing value replacing

* added test for new transform

* fix test

* Update src/mlnet/Templates/Console/MLCodeGen.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* Sanitize the column names in CLI (#162)

* added sanitization layer in CLI

* fix test

* changed exception.StackTrace to exception.ToString()

* fix package name (#168)

* Rev public API (#163)

* Rename TransformGeneratorBase .cs to TransformGeneratorBase.cs (#153)

* Fix minor version for the repository + remove Nlog config package (#171)

*  changed the minor version

* removed the nlog config package

* Added new test to columninfo and fixing up API (#178)

* Make optimizing metric customizable and add trainer whitelist functionality (#172)

* API rev (#181)

* propagate root MLContext thru AutoML (instead of creating our own) (#182)

* Enabling new command line args (#183)

* fix package name

* initial commit

* added more commandline args

* fixed tests

* added headers

* fix tests

* fix test

* rename 'AutoFitter' to 'Experiment' (#169)

* added tests (#187)

* rev InferColumns to accept ColumnInfo input param (#186)

* Implement argument --has-header and change usage of dataset (#194)

* added has header and fixed dataset and train dataset

* fix tests

* removed dummy command (#195)

* Fix bug for regression and sanitize input label from user (#198)

* removed dummy command

* sanitize label and fix template

* fix tests

* Do not generate code concatenating columns when the dataset has a single feature column (#191)

* Include some missed logging in the generated code.  (#199)

* added logging messages for generated code

* added log messages

* deleted file

* cleaning up proj files (#185)

* removed platform target

* removed platform target

* Some spaces and extra lines + bug in output path  (#204)

* nit picks

* nit picks

* fix test

* accept label from user input and provide in generated code (#205)

* Rev handling of weight / label columns (#203)

* migrate to private ML.NET nuget for latest bug fixes (#131)

* fix multiclass with nonstandard label (#207)

* Multiclass nondefault label test (#208)

* printing escaped chars + bug (#212)

* delete unused internal samples (#211)

* fix SMAC bug that causes multiclass sample to infinite loop (#209)

* Rev user input validation for new API (#210)

* added console message for exit and nit picks (#215)

* exit when exception encountered (#216)

* Seal API classes (and make EnableCaching internal) (#217)

* Suggested sample nits (feel free to ask for any of these to be reverted) (#219)

* User input column type validation (#218)

* upgrade commandline and renaming (#221)

* upgrade commandline and renaming

* renaming fields

* Make build.sh, init-tools.sh, & run.sh executable on OSX/Linux (#225)

*  CLI argument descriptions updated (#224)

* CLI argument descriptions updated

* No version in .csproj

* added flag to disable training code (#227)

* Exit if perfect model produced (#220)

* removed header (#228)

* removed header

* added auto generated header

* removed console read key (#229)

* Fix model path in generated file (#230)

* removed console read key

* fix model path

* fix test

* reorder samples (#231)

* remove rule that infers column purpose as categorical if # of distinct values is < 100 (#233)

* Null reference exception fix for finding best model when some runs have failed (#239)

* samples fixes (#238)

* fix for defaulting Averaged Perceptron # of iterations to 10 (#237)

* Bug bash feedback Feb 27. API changes and sample changes (#240)

* Bug bash feedback Feb 27. 
API changes 
Sample changes
Exception fix

* Samples / API rev from 2/27 bug bash feedback (#242)

* changed the directory structure for generated project (#243)

* changed the directory structure for generated project

* changed test

* upgraded commandline package

* Fix test file locations on OSX (#235)

* fix test file locations on OSX

* changing to Path.Combine()

* Additional Path.Combine()

* Remove ConsoleCodeGeneratorTests.GeneratedTrainCodeTest.received.txt

* Additional Path.Combine()

* add back in double comparison fix

* remove metrics agent NaN returns

* test fix

* test format fix

* mock out path

Thanks to @daholste for additional fixes!

* upgrade to latest ML.NET public surface (#246)

* Upgrade to ML.NET 0.11 (#247)

* initial changes

* fix lightgbm

* changed normalize method

* added tests

* fix tests

* fix test

* Private preview final API changes (#250)

* .NET framework design guidelines applied to public surface
* WhitelistedTrainers -> Trainers

* Add estimator to public API iteration result (#248)

* LightGBM pipeline serialization fix (#251)

* Change order that we search for TextLoader's parameters (#256)

* CLI IFileInfo null exception fix (#254)

* Averaged Perceptron pipeline serialization fix (#257)

* Upgrade command-line-api and default folder name change (#258)

* change in defautl folderName

* upgrade command line

* Update src/mlnet/Program.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* eliminate IFileInfo from CLI (#260)

* Rev samples towards private preview; ignored columns fix (#259)

* remove unused methods in consolehelper and nit picks in generated code (#261)

* nit picks

* change in console helper

* fix tests

* add space

* fix tests

* added nuget sources in generated csproj (#262)

* added nuget sources in csproj

* changed the structure in generated code

* space

* upgrade to mlnet 0.11 (#263)

* Formatting CLI metrics (#264)

Ensures space between printed metrics (also model counter). Right aligned metrics. Extended AUC to four digits.

* Add implementation of non -ova multi class trainers code gen (#267)

* added non ova multi class learners

* added tests

* test cases

* Add caching (#249)

* AdvancedExperimentSettings sample nits (#265)

* Add sampling key column (#268)

* Initial work for multi-class classification support for CLI (#226)

* Initial work for multi-class classification support for CLI

* String updates

* more strings

* Whitelist non-OVA multi-class learners

* Refactor the orchestration of AutoML calls (#272)

* Do not auto-group columns with suggested purpose = 'Ignore' (#273)

* Fix: during type inferencing, parse whitespace strings as NaN (#271)

* Printing additional metrics in CLI for binary classification (#274)

* Printing additional metrics in CLI for binary classification

* Update src/mlnet/Utilities/ConsolePrinter.cs

* Add API option to store models on disk (instead of in memory); fix IEstimator memory leak (#269)

* Print failed iterations in CLI (#275)

* change the type to float from double (#277)

* cache arg implementation in CLI (#280)

* cache implementation

* corrected the null case

* added tests for all cases

* Remove duplicate value-to-key mapping transform for multiclass string labels (#283)

* Add post-trainer transform SDK infra; add KeyToValueMapping transform to CLI; fix: for generated multiclass models, convert predicted label from key to original label column type (#286)

* Implement ignore columns command line arg (#290)

* normalize line endings

* added --ignore-columns

* null checks

* unit tests

* Print winning iteration and runtime in CLI (#288)

* Print best metric and runtime

* Print best metric and runtime

* Line endings in AutoMLEngine.cs

* Rename time column to duration to match Python SDK

* Revert to MicroAccuracy and MacroAccuracy spellings

* Revert spelling of BinaryClassificationMetricsAgent to BinaryMetricsAgent to reduce merge conflicts

* Revert spelling of MulticlassMetricsAgent to MultiMetricsAgent to reduce merge conflicts

* missed some files

* Fix merge conflict

* Update AutoMLEngine.cs

* Add MacOS & Linux to CI; MacOS & Linux test fixes (#293)

* MicroAccuracy as default for multi-class (#295)

Change default optimization metric for multi-class classification to MicroAccuracy (accuracy). Previously it was set to MacroAccuracy.

* Null exception for ignorecolumns in CLI (#294)

* Null exception for ignorecolumns in CLI

* Check if ignore-columns array has values (as the default is now a empty array)

* Emit caching flag in pipeline object model. (Includes SuggestedPipelineBuilder refactor & debug string fixes / refactor) (#296)

* removed sln (#297)

* Caching enabling in code gen part -2 (#298)

* add

* added caching codegen

* support comma separated values for --ignore-columns (#300)

* default initialization for ignore columns (#302)

* default initialization

* adde null check

* Codegen for multiclass non-ova (#303)

* changes to template

* multicalss codegen

* test cases

* fix test cases

* Generated Project new structure. (#305)

* added new templates

* writing files to disck

* change path

* added new templates

* misisng braces

* fix bugs

* format code

* added util methods for solution file creation and addition of projects to it

* added extra packages to project files

* new tests

* added correct path for sln

* build fix

* fix build

* include using system in prediction class (#307)

* added using

* fix test

* Random number generator is not thread safe (#310)

* Random number generator is not thread safe

* Another local random generator

* Missed a few references

* Referncing AutoMlUtils.random instead of a local RNG

* More refs to mail RNG; remove Float as per https://github.com/dotnet/machinelearning/issues/1669

* Missed Random.cs

* Fix multiclass code gen (#314)

* compile error in codegen

* removes scores printing

* fix bugs

* fix test

* Fix compile error in codegen project (#319)

* removed redundant code

* fix test case

* Rev OVA pipeline node SDK output: wrap binary trainers as children inside parent OVA node (#317)

* Ova Multi class codegen support (#321)

* dummy

* multiova implementation

* fix tests

* remove inclusion list

* fix tests and console helper

* Rev run result trainer name for OVA: output different trainer name for each OVA + binary learner combination (#322)

* Rev run result trainer name for Ova: output different trainer name for each Ova + binary learner combination

* test fixes

* Console helper bug in generated code for multiclass (#323)

* fix

* fix test

* looping perlogclass

* fix test

* Initial version of Progress bar impl and CLI UI experience (#325)

* progressbar

* added progressbar and refactoring

* reverted

* revert sign assembly

* added headers and removed exception rethrow

* Setting model directory to temp directory (#327)

* Suggested changes to progress bar (#335)

* progressbar

* added progressbar and refactoring

* reverted

* revert sign assembly

* added headers and removed exception rethrow

* bug fixes and updates to UI

* added friendly name printing for metric

* formatting

* Rev Samples (#334)

* Telemetry2 (#333)

* Create test.txt

* Create test.txt

* changes needed for benchmarking

* forgot one file

* merge conflict fix

* fix build break

* back out my version of the fix for Label column issue and fix the original fix

* bogus file removal

* undo SuggestedPipeline change

* remove labelCol from pipeline suggester

* fix build break

* rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline)

* tweak queue in vsts-ci.yml

* CLI telemetry implementation

* Telemetry implementation

* delete unnecessary file and change file size bucket to actually log log2 instead of nearest ceil value

* add headers, remove comments

* one more header missing

* Fix progress bar in linux/osx (#336)

* progressbar

* added progressbar and refactoring

* reverted

* revert sign assembly

* added headers and removed exception rethrow

* bug fixes and updates to UI

* added friendly name printing for metric

* formatting

* change from task to thread

* Update src/mlnet/CodeGenerator/CodeGenerationHelper.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* Mem leak fix (#328)

* Create test.txt

* Create test.txt

* changes needed for benchmarking

* forgot one file

* merge conflict fix

* fix build break

* back out my version of the fix for Label column issue and fix the original fix

* bogus file removal

* undo SuggestedPipeline change

* remove labelCol from pipeline suggester

* fix build break

* rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline)

* tweak queue in vsts-ci.yml

* there is still investigation to be done but this fix works and solves memory leak problems

* minor refactor

* Upgrade ML.NET package (#343)

* Add cross-validation (CV), and auto-CV for small datasets; push common API experiment methods into base class (#287)

* restore old yml for internal pipeline so we can publish nuget again to devdiv stream (#344)

* Polishing the CLI UI part-1 (#338)

* formatting of pbar message

* Polishing the UI

* optimization

* rename variable

* Update src/mlnet/AutoML/AutoMLEngine.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* Update src/mlnet/CodeGenerator/CodeGenerationHelper.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* new message

* changed hhtp to https

* added iteration num + 1

* change string name and add color to artifacts

* change the message

* build errors

* added null checks

* added exception messsages to log file

* added exception messsages to log file

* CLI ML.NET version upgrade (#345)

* Sample revs; ColumnInformation property name revs; pre-featurizer fixes (#346)

* CLI -- consume logs from AutoML SDK (#349)

* Rename RunDetails --> RunDetail (#350)

* command line api upgrade and progress bar rendering bug (#366)

* added fix for all platforms progress bar

* upgrade nuget

* removed args from writeline

* change in the version (#368)

* fix few bugs in progressbar and verbosity (#374)

* fix few bugs in progressbar and verbosity

* removed unused name space

* Fix for folders with space in it while generating project (#376)

* support for folders with spaces

* added support for paths with space

* revert file

* change name of var

* remove spaces

* SMAC fix for minimizing metrics (#363)

* Formatting Regression metrics and progress bar display days. (#379)

* added progress bar day display and fix regression metrics

* fix formatting

* added total time

* formatted total time

* change command name and add pbar message (#380)

* change command name and add pbar message

* fix tests

* added aliases

* duplicate alias

* added another alias for task

* UI missing features (#382)

* added formatting changes

* added accuracy specifically

* downgrade the codepages (#384)

* Change in project structure (#385)

* initial changes

* Change in project structure

* correcting test

* change variable name

* fix tests

* fix tests

* fix more tests

* fix codegen errors

* adde log file message

* changed name of args

* change variable names

* fix test

* FileSizeBuckets in correct units (#387)

* Minor telemetry change to log in correct units and make our life easier in the future

* Use Ceiling instead of Round

* changed order (#388)

* prep work to transfer to ml.net (#389)

* move test projects to top level test subdir

* rename some projects to make naming consistent and make it build again

* fix test project refs

* Add AutoML components to build, fix issues related to that so it builds

* fix test cases, remove AppInsights ref from AutoML (#3329)

* [AutoML] disable netfx build leg for now (#3331)

* disable netfx build leg for now

* disable netfx build leg for now.

* [AutoML] Add AutoML XML documentation to all public members; migrate AutoML projects & tests into ML.NET solution; AutoML test fixes (#3351)

* [AutoML] Rev AutoML public API; add required native references to AutoML projects (#3364)

* [AutoML] Minor changes to generated project in CLI based on feedback (#3371)

* nitpicks for generated project

* revert back the target framework

* [AutoML] Migrate AutoML back to its own solution, w/ NuGet dependencies (#3373)

* Migrate AutoML back to its own solution, w/ NuGet dependencies

* build project updates; parameter name revert

* dummy change

* Revert "dummy change"

This reverts commit 3e8574266f556a4d5b6805eb55b4d8b8b84cf355.

* [AutoML] publish AutoML package (#3383)

* publish AutoML package

* Only leave automl and mlnet tests to run

* publish AutoML package

* Only leave automl and mlnet tests to run

* fix build issues when ml.net is not building

* bump version to 0.3 since that's the one we're going to ship for build (#3416)

* [AutoML] temporarily disable all but x64 platforms -- don't want to do native builds and can't find a way around that with the current VSTS pipeline (#3420)

* disable steps but keep phases to keep vsts build pipeline happy (#3423)

* API docs for experimentation (#3484)

* fixed path bug and regression metrics correction (#3504)

* changed the casing of option alias as it conflicts with --help (#3554)

* [AutoML] Generated project - FastTree nuget package inclusion dynamically (#3567)

* added support for fast tree nuget pack inclusion in generated project

* fix testcase

* changed the tool name in telemetry message

* dummy commit

* remove space

* dummy commit to trigger build

* [AutoML] Add AutoML example code (#3458)

* AutoML PipelineSuggester: don't recommend pipelines from first-stage trainers that failed (#3593)

* InferColumns API: Validate all columns specified in column info exist in inferred data view (#3599)

* [AutoML] AutoML SDK API: validate schema types of input IDataView (#3597)

* [AutoML] If first three iterations all fail, short-circuit AutoML experiment (#3591)

* mlnet CLI nupkg creation/signing (#3606)

* mlnet CLI nupkg creation/signing

* relmove includeinpackage from mlnet csproj

* address PR comments -- some minor reshuffling of stuff

* publish symbols for mlnet CLI

* fix case in NLog.config

* [AutoML] rename Auto to AutoML in namespace and nuget (#3609)

* mlnet CLI nupkg creation/signing

* [AutoML] take dependency on a specific ml.net version (#3610)

* take dependency on a specific ml.net version

* catch up to spelling fix for OptimizationTolerance

* force a specific ml.net nuget version, fix typo (#3616)

* [AutoML] Fix error handling in CLI.  (#3618)

* fix error handling

* renaming variables

* [AutoML] turn off line pragmas in .tt files to play nice with signing (#3617)

* turn off line pragmas in .tt files to play nice with signing

* dedupe tags

* change the param name (#3619)

* [AutoML]  return null instead of null ref crash on Model property accessor (#3620)

* return null instead of null ref crash on Model property accessor

* [AutoML] Handling label column names which have space and exception logging (#3624)

* fix case of label with space and exception logging

* final handler

* revert file

* use Name instead of FullName for telemetry filename hash (#3633)

* renamed classes (#3634)

* change ML.NET dependency to 1.0 (#3639)

[AutoML] undo pinning ML.NET dependency

* set exploration time default in CLI to half hour (#3640)

* [AutoML] step 2 of removing pinned nupkg versions (#3642)

* InferColumns API that consumes label column index -- Only rename label column to 'Label' for headerless files (#3643)

* [AutoML] Upgrade ml.net package in generated code (#3644)

* upgrade the mlnet package in gen code

* Update src/mlnet/Templates/Console/ModelProject.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* Update src/mlnet/Templates/Console/ModelProject.tt

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* added spaces

* [AutoML] Early stopping in CLI based on the exploration time (#3641)

* early stopping in CLI

* remove unused variables

* change back to thread

* remove sleep

* fix review comments

* remove ununsed usings

* format message

* collapse declaration

* remove unused param

* added environment.exit and removal of error message

* correction in message

* secs-> seconds

* exit code

* change value to 1

* reverse the declaration

* [AutoML] Change wording for CouldNotFinshOnTime message (#3655)

* set exploration time default in CLI to half hour

* [AutoML] Change wording for CouldNotFinshOnTime message

* [AutoML] Change wording for CouldNotFinshOnTime message

* even better wording for CouldNotFinshOnTime

* temp change to get around vsts publish failure (#3656)

* [AutoML] bump version to 0.4.0 (#3658)

* implement culture invariant strings (#3725)

* reset culture (#3730)

* [AutoML] Cross validation fixes; validate empty training / validation input data (#3794)

* [AutoML] Enable style cop rules & resolve errors (#3823)

* add task agnostic wrappers for autofit calls (#3860)

* [AutoML] CLI telemetry rev (#3789)

* delete automl .sln

* CLI -- regenerate templated CS files (#3954)

* [AutoML] Bump ML.NET package version to 1.2.0 in AutoML API and CLI; and AutoML package versions to 0.14.0 (#3958)

* Build AutoML NuGet package (#3961)

* Increment AutoML build version to 0.15.0 for preview. (#3968)

* added culture independent parsing (#3731)

* - convert tests to xunit
- take project level dependency on ML.NET components instead of nuget
- set up bestfriends relationship to ML.Core and remove some of the copies of util classes from AutoML.NET (more work needed to fully remove them, work item 4064)
- misc build script changes to address PR comments

* address issues only showing up in a couple configurations during CI build

* fix cut&paste error

* [AutoML] Bump version to ML.NET 1.3.1 in AutoML API and CLI and AutoML package version to 0.15.1 (#4071)

* bumped version

* change versions in nupkg

* revert version bump in branch props

* [AutoML] Fix for Exception thrown in cross val when one of the score equals infinity. (#4073)

* bumped version

* change versions in nupkg

* revert version bump in branch props

* added infinity fix

* changes signing (#4079)

* Addressed PR comments and build issues
- sync block on creating test data file (failed intermittently)
- removed classes we copied over from ML.Core and fixed their uses to de-dupe and use original ML.Core versions since we now have InternalsVisible and BestFriends
- Fixed nupkg creation  to use projects insted of public nuget version for AutoML
- Fixed a bunch of unit tests that didn't actually test what they were supposed to test, while removing cut&past code and dependencies.
- Few more misc small changes

* minor nit - removed unused folder ref

* Fix the .sln file for the right configurations.

* Fix mistake in .sln file

* test fixes and disable one test

* fix tests, re-add AutoML samples csproj

* bumped VS version to 16 in .sln, removed InternalsVisible for a dead assembly, removed unused references from AutoML test project

* Updated docs to include PredictedLabel member (#4107)

* Fixed build errors resulting from upgrade to VS2019 compilers

* Added additional message describing the previous fix

* Updated docs to include PredictedLabel member

* Added CODEOWNERS file in the .github/ folder. (#4140)

* Added CODEOWNERS file in the .github/ folder. This allows reviewers to review any changes in the machine learning repository

* Updated .github/CODEOWNERS with the team instead of individual reviewers

* Added AutoML team reviewers (#4144)

* Added CODEOWNERS file in the .github/ folder. This allows reviewers to review any changes in the machine learning repository

* Updated .github/CODEOWNERS with the team instead of individual reviewers

* Added AutoML team reviwers to files owned by AutoML team

* Added AutoML team reviwers to files owned by AutoML team

* Removed two files that don't exist for AutoML team in CODEOWNERS

* Build extension method to reload changes without specifying model name (#4146)

* Image classification preview 2. (#4151)

* Image classification preview 2.

* PR feedback.

* Add unit-test.

* Add unit-test.

* Add unit-test.

* Add unit-test.

* Use Path.Combine instead of Join.

* fix test dataset path.

* fix test dataset path.

* Improve test.

* Improve test.

* Increase epochs in tests.

* Disable test on Ubuntu.

* Move test to its own project.

* Move test to its own project.

* Move test to its own project.

* Move test to its own file.

* cleanup.

* Disable parallel execution of tensorflow tests.

* PR feedback.

* PR feedback.

* PR feedback.

* PR feedback.

* Prevent TF test to execute in parallel.

* PR feedback.

* Build error.

* clean up.

* Syncing upstream fork (#11)

* Throw error on incorrect Label name in InferColumns API (#47)

* Added sequential grouping of columns

* reverted the file

* addded infer columns label name checking

* added column detection error

* removed unsed usings

* added quotes

* replace Where with Any clause

* replace Where with Any clause

* Set Nullable Auto params to null values (#50)

* Added sequential grouping of columns

* reverted the file

* added auto params as null

* change to the update fields method

* First public api propsal (#52)

* Includes following
1) Final proposal for 0.1 public API surface
2) Prefeaturization
3) Splitting train data into train and validate when validation data is null
4) Providing end to end samples one each for regression, binaryclassification and multiclass classification

* Incorporating code review feedbacks

* Revert "Set Nullable Auto params to null values" (#53)

* Revert "First public api propsal (#52)"

This reverts commit e4a64cf4aeab13ee9e5bf0efe242da3270241bd7.

* Revert "Set Nullable Auto params to null values (#50)"

This reverts commit 41c663cd14247d44022f40cf2dce5977dbab282d.

* AutoFit return type is now an IEnumerable (#55)

AutoFit returns is now an IEnumerable - this enables many good things

Implementing variety of early stopping criteria (See sample)
Early discard of models that are no good. This improves memory usage efficiency. (See sample)
No need to implement a callback to get results back
Getting best score is now outside of API implementation. It is a simple math function to compare scores (See sample).

Also templatized the return type for better type safety through out the code.

* misc fixes & test additions, towards 0.1 release (#56)

* Enable UnitTests on build server (#57)

* 1) Making trainer name public (#62)

2) Fixing up samples to reflect it

*  Initial version of CLI tool for mlnet (#61)

* added global tool initial project

* removed unneccesary files, renamed files

* refactoring and added base abstract classes for trainer generator

* removed unused class

* Added classes for transforms

* added transform generate dummy classes

* more refactoring, added first transform

* more refactoring and added classes

* changed the project structure

* restructing added options class

* sln changes

* refactored options to different class:

* added more logic for code generation of class

* misc changes

* reverted file

* added commandline api package

* reverted sample

* added new command line api parser

* added normalization of column names

* Added command defaults and error message

* implementation of all trainers

* changed auto to null

* added all transform generators

* added error handling when args is empty and minor changes due to change in AutoML api names

* changed the name of param

* added new command line options and restructuring code

* renamed proj file and added solution

* Added code to generate usings, Fixed few bugs in the code

* added validation to the command line options

* changed project name

* Bug fixes due to API change in AutoML

* changed directory structure

* added test framework and basic tests

* added more tests

* added improvements to template and error handling

* renamed the estimator name

* fixed test case

* added comments

* added headers

* changed namespace and removed unneccesary properties from project

* Revert "changed namespace and removed unneccesary properties from project"

This reverts commit 9edae033e9845e910f663f296e168f1182b84f5f.

* fixed test cases and renamed namespaces

* cleaned up proj file

* added folder structure

* added symbols/tokens for strings

* added more tests

* review comments

* modified test cases

* review comments

* change in the exception message

* normalized line endings

* made method private static

* simplified range building /optimization

* minor fix

* added header

* added static methods in command where necessary

* nit picks

*  made few methods static

* review comments

* nitpick

* remove line pragmas

* fix test case

* Use better AutiFit overload and ignore Multiclass (#64)

* Upgrading CLI to produce ML.NET V.10 APIs and bunch of Refactoring tasks (#65)

* Added sequential grouping of columns

* reverted the file

* upgrade to v .10 and refactoring

* added null check

* fixed unit tests

* review comments

* removed the settings change

* added regions

* fixed unit tests

* Upgrade ML.NET package to 0.10.0 (#70)

* Change in template to accomodate new API of TextLoader (#72)

* Added sequential grouping of columns

* reverted the file

* changed to new API of Text Loader

* changed signature

* added params for taking additional settings

* changes to codegen params

* refactoring of templates and fixing errors

* Enable gated check for mlnet.tests (#79)

* Added sequential grouping of columns

* reverted the file

* changed to new API of Text Loader

* changed signature

* added params for taking additional settings

* changes to codegen params

* refactoring of templates and fixing errors

* added run-tests.proj and referred it in build.proj

* CLI tool - make validation dataset optional and support for crossvalidation in generated code (#83)

* Added sequential grouping of columns

* reverted the file

* bug fixes, more logic to templates to support cross-validate

* formatting and fix type in consolehelper

* Added logic in templates

* revert settings

* benchmarking related changes (#63)

* Create test.txt

* Create test.txt

* changes needed for benchmarking

* forgot one file

* merge conflict fix

* fix build break

* back out my version of the fix for Label column issue and fix the original fix

* bogus file removal

* undo SuggestedPipeline change

* remove labelCol from pipeline suggester

* fix build break

* fix fast forest learner (don't sweep over learning rate) (#88)

* Made changes to Have non-calibrated scoring for binary classifiers (#86)

* Added sequential grouping of columns

* reverted the file

* added calibration workaround

* removed print probability

* reverted settings

* rev ColumnInference API: can take label index; rev output object types; add tests (#89)

* rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (#99)

* Create test.txt

* Create test.txt

* changes needed for benchmarking

* forgot one file

* merge conflict fix

* fix build break

* back out my version of the fix for Label column issue and fix the original fix

* bogus file removal

* undo SuggestedPipeline change

* remove labelCol from pipeline suggester

* fix build break

* rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline)

* publish nuget (#101)

* use dotnet-internal-temp agent for internal build

* use dotnet-internal feed

* Fix Codegen for columnConvert and ValueToKeyMapping transform and add individual transform tests (#95)

* Added sequential grouping of columns

* reverted the file

* fix usings for type convert

* added transforms tests

* review comments

* When generating usings choose only distinct usings directives (#94)

* Added sequential grouping of columns

* reverted the file

* Added code to have unique strings

* refactoring

* minor fix

* minor fix

* Autofit overloads + cancellation + progress callbacks

1) Introduce AutoFit overloads (basic and advanced)
2) AutoFit Cancellation
3) AutoFit progress callbacks

* Default the kfolds to value 5 in CLI generated code (#115)

* Added sequential grouping of columns

* reverted the file

* Set up CI with Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* remove file

* added kfold param and defaulted to value

* changed type

* added for regression

* Remove extra ; from generated code (#114)

* Added sequential grouping of columns

* reverted the file

* Set up CI with Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* removed extra ; from generated code

* removed file

* fix unit tests

* TimeoutInSeconds (#116)

Specifying timeout in seconds instead of minutes

* Added more command line args implementation to CLI tool and refactoring (#110)

* Added sequential grouping of columns

* reverted the file

* Set up CI with Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* Update azure-pipelines.yml for Azure Pipelines

* added git status

* reverted change

* added codegen options and refactoring

* minor fixes'

* renamed params, minor refactoring

* added tests for commandline and refactoring

* removed file

* added back the test case

* minor fixes

* Update src/mlnet.Test/CommandLineTests.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* review comments

*  capitalize the first character

* changed the name of test case

* remove unused directives

* Fail gracefully if unable to instantiate data view with swept parameters (#125)

* gracefully fail if fail to parse a datai

* rev

* validate AutoFit 'Features' column must be of type R4 (#132)

* Samples: exceptions / nits (#124)

* Logging support in CLI + Implementation of cmd args [--name,--output,--verbosity] (#121)

* addded logging and helper methods

* fixing code after merge

* added resx files, added logger framework, added logging messages

* added new options

* added spacing

* minor fixes

* change command description

* rename option, add headers, include new param in test

* formatted

* build fix

*  changed option name

* Added NlogConfig file

* added back config package

* fix tests

* added correct validation check (#137)

* Use CreateTextLoader<T>(..)  instead of CreateTextLoader(..) (#138)

* added support to loaddata by class in the generated code

* fix tests

* changed CreateTextLoader to ReadFromTextFile method. (#140)

* changed textloader to readfromtextfile method

* formatting

* exception fixes (#136)

* infer purpose of hidden columns as 'ignore' (#142)

* Added approval tests and bunch of refactoring of code and normalizing namespaces (#148)

* changed textloader to readfromtextfile method

* formatting

* added approval tests and refactoring of code

* removed few comments

* API 2.0 skeleton (#149)

Incorporating API review feedback

* The CV code should come before the training when there is no test dataset in generated code (#151)

* reorder cv code

* build fix

* fixed structure

* Format the generated code + bunch of misc tasks (#152)

* added formatting and minor changes for reordering cv

* fixing the template

* minor changes

* formatting changes

* fixed approval test

* removed unused nuget

* added missing value replacing

* added test for new transform

* fix test

* Update src/mlnet/Templates/Console/MLCodeGen.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* Sanitize the column names in CLI (#162)

* added sanitization layer in CLI

* fix test

* changed exception.StackTrace to exception.ToString()

* fix package name (#168)

* Rev public API (#163)

* Rename TransformGeneratorBase .cs to TransformGeneratorBase.cs (#153)

* Fix minor version for the repository + remove Nlog config package (#171)

*  changed the minor version

* removed the nlog config package

* Added new test to columninfo and fixing up API (#178)

* Make optimizing metric customizable and add trainer whitelist functionality (#172)

* API rev (#181)

* propagate root MLContext thru AutoML (instead of creating our own) (#182)

* Enabling new command line args (#183)

* fix package name

* initial commit

* added more commandline args

* fixed tests

* added headers

* fix tests

* fix test

* rename 'AutoFitter' to 'Experiment' (#169)

* added tests (#187)

* rev InferColumns to accept ColumnInfo input param (#186)

* Implement argument --has-header and change usage of dataset (#194)

* added has header and fixed dataset and train dataset

* fix tests

* removed dummy command (#195)

* Fix bug for regression and sanitize input label from user (#198)

* removed dummy command

* sanitize label and fix template

* fix tests

* Do not generate code concatenating columns when the dataset has a single feature column (#191)

* Include some missed logging in the generated code.  (#199)

* added logging messages for generated code

* added log messages

* deleted file

* cleaning up proj files (#185)

* removed platform target

* removed platform target

* Some spaces and extra lines + bug in output path  (#204)

* nit picks

* nit picks

* fix test

* accept label from user input and provide in generated code (#205)

* Rev handling of weight / label columns (#203)

* migrate to private ML.NET nuget for latest bug fixes (#131)

* fix multiclass with nonstandard label (#207)

* Multiclass nondefault label test (#208)

* printing escaped chars + bug (#212)

* delete unused internal samples (#211)

* fix SMAC bug that causes multiclass sample to infinite loop (#209)

* Rev user input validation for new API (#210)

* added console message for exit and nit picks (#215)

* exit when exception encountered (#216)

* Seal API classes (and make EnableCaching internal) (#217)

* Suggested sample nits (feel free to ask for any of these to be reverted) (#219)

* User input column type validation (#218)

* upgrade commandline and renaming (#221)

* upgrade commandline and renaming

* renaming fields

* Make build.sh, init-tools.sh, & run.sh executable on OSX/Linux (#225)

*  CLI argument descriptions updated (#224)

* CLI argument descriptions updated

* No version in .csproj

* added flag to disable training code (#227)

* Exit if perfect model produced (#220)

* removed header (#228)

* removed header

* added auto generated header

* removed console read key (#229)

* Fix model path in generated file (#230)

* removed console read key

* fix model path

* fix test

* reorder samples (#231)

* remove rule that infers column purpose as categorical if # of distinct values is < 100 (#233)

* Null reference exception fix for finding best model when some runs have failed (#239)

* samples fixes (#238)

* fix for defaulting Averaged Perceptron # of iterations to 10 (#237)

* Bug bash feedback Feb 27. API changes and sample changes (#240)

* Bug bash feedback Feb 27. 
API changes 
Sample changes
Exception fix

* Samples / API rev from 2/27 bug bash feedback (#242)

* changed the directory structure for generated project (#243)

* changed the directory structure for generated project

* changed test

* upgraded commandline package

* Fix test file locations on OSX (#235)

* fix test file locations on OSX

* changing to Path.Combine()

* Additional Path.Combine()

* Remove ConsoleCodeGeneratorTests.GeneratedTrainCodeTest.received.txt

* Additional Path.Combine()

* add back in double comparison fix

* remove metrics agent NaN returns

* test fix

* test format fix

* mock out path

Thanks to @daholste for additional fixes!

* upgrade to latest ML.NET public surface (#246)

* Upgrade to ML.NET 0.11 (#247)

* initial changes

* fix lightgbm

* changed normalize method

* added tests

* fix tests

* fix test

* Private preview final API changes (#250)

* .NET framework design guidelines applied to public surface
* WhitelistedTrainers -> Trainers

* Add estimator to public API iteration result (#248)

* LightGBM pipeline serialization fix (#251)

* Change order that we search for TextLoader's parameters (#256)

* CLI IFileInfo null exception fix (#254)

* Averaged Perceptron pipeline serialization fix (#257)

* Upgrade command-line-api and default folder name change (#258)

* change in defautl folderName

* upgrade command line

* Update src/mlnet/Program.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* eliminate IFileInfo from CLI (#260)

* Rev samples towards private preview; ignored columns fix (#259)

* remove unused methods in consolehelper and nit picks in generated code (#261)

* nit picks

* change in console helper

* fix tests

* add space

* fix tests

* added nuget sources in generated csproj (#262)

* added nuget sources in csproj

* changed the structure in generated code

* space

* upgrade to mlnet 0.11 (#263)

* Formatting CLI metrics (#264)

Ensures space between printed metrics (also model counter). Right aligned metrics. Extended AUC to four digits.

* Add implementation of non -ova multi class trainers code gen (#267)

* added non ova multi class learners

* added tests

* test cases

* Add caching (#249)

* AdvancedExperimentSettings sample nits (#265)

* Add sampling key column (#268)

* Initial work for multi-class classification support for CLI (#226)

* Initial work for multi-class classification support for CLI

* String updates

* more strings

* Whitelist non-OVA multi-class learners

* Refactor the orchestration of AutoML calls (#272)

* Do not auto-group columns with suggested purpose = 'Ignore' (#273)

* Fix: during type inferencing, parse whitespace strings as NaN (#271)

* Printing additional metrics in CLI for binary classification (#274)

* Printing additional metrics in CLI for binary classification

* Update src/mlnet/Utilities/ConsolePrinter.cs

* Add API option to store models on disk (instead of in memory); fix IEstimator memory leak (#269)

* Print failed iterations in CLI (#275)

* change the type to float from double (#277)

* cache arg implementation in CLI (#280)

* cache implementation

* corrected the null case

* added tests for all cases

* Remove duplicate value-to-key mapping transform for multiclass string labels (#283)

* Add post-trainer transform SDK infra; add KeyToValueMapping transform to CLI; fix: for generated multiclass models, convert predicted label from key to original label column type (#286)

* Implement ignore columns command line arg (#290)

* normalize line endings

* added --ignore-columns

* null checks

* unit tests

* Print winning iteration and runtime in CLI (#288)

* Print best metric and runtime

* Print best metric and runtime

* Line endings in AutoMLEngine.cs

* Rename time column to duration to match Python SDK

* Revert to MicroAccuracy and MacroAccuracy spellings

* Revert spelling of BinaryClassificationMetricsAgent to BinaryMetricsAgent to reduce merge conflicts

* Revert spelling of MulticlassMetricsAgent to MultiMetricsAgent to reduce merge conflicts

* missed some files

* Fix merge conflict

* Update AutoMLEngine.cs

* Add MacOS & Linux to CI; MacOS & Linux test fixes (#293)

* MicroAccuracy as default for multi-class (#295)

Change default optimization metric for multi-class classification to MicroAccuracy (accuracy). Previously it was set to MacroAccuracy.

* Null exception for ignorecolumns in CLI (#294)

* Null exception for ignorecolumns in CLI

* Check if ignore-columns array has values (as the default is now a empty array)

* Emit caching flag in pipeline object model. (Includes SuggestedPipelineBuilder refactor & debug string fixes / refactor) (#296)

* removed sln (#297)

* Caching enabling in code gen part -2 (#298)

* add

* added caching codegen

* support comma separated values for --ignore-columns (#300)

* default initialization for ignore columns (#302)

* default initialization

* adde null check

* Codegen for multiclass non-ova (#303)

* changes to template

* multicalss codegen

* test cases

* fix test cases

* Generated Project new structure. (#305)

* added new templates

* writing files to disck

* change path

* added new templates

* misisng braces

* fix bugs

* format code

* added util methods for solution file creation and addition of projects to it

* added extra packages to project files

* new tests

* added correct path for sln

* build fix

* fix build

* include using system in prediction class (#307)

* added using

* fix test

* Random number generator is not thread safe (#310)

* Random number generator is not thread safe

* Another local random generator

* Missed a few references

* Referncing AutoMlUtils.random instead of a local RNG

* More refs to mail RNG; remove Float as per https://github.com/dotnet/machinelearning/issues/1669

* Missed Random.cs

* Fix multiclass code gen (#314)

* compile error in codegen

* removes scores printing

* fix bugs

* fix test

* Fix compile error in codegen project (#319)

* removed redundant code

* fix test case

* Rev OVA pipeline node SDK output: wrap binary trainers as children inside parent OVA node (#317)

* Ova Multi class codegen support (#321)

* dummy

* multiova implementation

* fix tests

* remove inclusion list

* fix tests and console helper

* Rev run result trainer name for OVA: output different trainer name for each OVA + binary learner combination (#322)

* Rev run result trainer name for Ova: output different trainer name for each Ova + binary learner combination

* test fixes

* Console helper bug in generated code for multiclass (#323)

* fix

* fix test

* looping perlogclass

* fix test

* Initial version of Progress bar impl and CLI UI experience (#325)

* progressbar

* added progressbar and refactoring

* reverted

* revert sign assembly

* added headers and removed exception rethrow

* Setting model directory to temp directory (#327)

* Suggested changes to progress bar (#335)

* progressbar

* added progressbar and refactoring

* reverted

* revert sign assembly

* added headers and removed exception rethrow

* bug fixes and updates to UI

* added friendly name printing for metric

* formatting

* Rev Samples (#334)

* Telemetry2 (#333)

* Create test.txt

* Create test.txt

* changes needed for benchmarking

* forgot one file

* merge conflict fix

* fix build break

* back out my version of the fix for Label column issue and fix the original fix

* bogus file removal

* undo SuggestedPipeline change

* remove labelCol from pipeline suggester

* fix build break

* rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline)

* tweak queue in vsts-ci.yml

* CLI telemetry implementation

* Telemetry implementation

* delete unnecessary file and change file size bucket to actually log log2 instead of nearest ceil value

* add headers, remove comments

* one more header missing

* Fix progress bar in linux/osx (#336)

* progressbar

* added progressbar and refactoring

* reverted

* revert sign assembly

* added headers and removed exception rethrow

* bug fixes and updates to UI

* added friendly name printing for metric

* formatting

* change from task to thread

* Update src/mlnet/CodeGenerator/CodeGenerationHelper.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* Mem leak fix (#328)

* Create test.txt

* Create test.txt

* changes needed for benchmarking

* forgot one file

* merge conflict fix

* fix build break

* back out my version of the fix for Label column issue and fix the original fix

* bogus file removal

* undo SuggestedPipeline change

* remove labelCol from pipeline suggester

* fix build break

* rename AutoML to Microsoft.ML.Auto everywhere and a shot at publishing nuget package (will probably need tweaks once I try to use the pipleline)

* tweak queue in vsts-ci.yml

* there is still investigation to be done but this fix works and solves memory leak problems

* minor refactor

* Upgrade ML.NET package (#343)

* Add cross-validation (CV), and auto-CV for small datasets; push common API experiment methods into base class (#287)

* restore old yml for internal pipeline so we can publish nuget again to devdiv stream (#344)

* Polishing the CLI UI part-1 (#338)

* formatting of pbar message

* Polishing the UI

* optimization

* rename variable

* Update src/mlnet/AutoML/AutoMLEngine.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* Update src/mlnet/CodeGenerator/CodeGenerationHelper.cs

Co-Authored-By: srsaggam <41802116+srsaggam@users.noreply.github.com>

* new message

* changed hhtp to https

* added iteration num + 1

* change string name and add color to artifacts

* change the message

* build errors

* added null checks

* added exception messsages to log file

* added exception messsages to log file

* CLI ML.NET version upgrade (#345)

* Sample revs; ColumnInformation property name revs; pre-featurizer fixes (#346)

* CLI -- consume logs from AutoML SDK (#349)

* Rename RunDetails --> RunDetail (#350)

* command line api upgrade and progress bar rendering bug (#366)

* added fix for all platforms progress bar

* upgrade nuget

* removed args from writeline

* change in the version (#368)

* fix few bugs in progressbar and verbosity (#374)

* fix few bugs in progressbar and verbosity

* removed unused name space

* Fix for folders with space in it while generating project (#376)

* support for folders with spaces

* added support for paths with space

* revert file

* change name of var

* remove spaces

* SMAC fix for minimizing metrics (#363)

* Formatting Regression metrics and progress bar display days. (#379)

* added progress bar day display and fix regression metrics

* fix formatting

* added total time

* formatted total time

* change command name and add pbar message (#380)

* change command name and add pbar message

* fix tests

* added aliases

* duplicate alias

* added another alias for task

* UI missing features (#382)

* added formatting changes

* added accuracy specifically

* downgrade the codepages (#384)

* Change in project structure (#385)

* initial changes

* Change in project structure

* correcting test

* change variable name

* fix tests

* fix tests

* fix more tests

* fix codegen errors

* adde log file message

* changed name of args

* change variable names

* fix test

* FileSizeBuckets in correct units (#387)

* Minor telemetry change to log in correct units and make our life easier in the future

* Use Ceiling instead of Round

* changed order (#388)

* prep work to transfer to ml.net (#389)

* move test projects to top level test subdir

* rename some projects to make naming consistent and make it build again

* fix test project refs

* Add AutoML components to build, fix issues related to that so it builds

* fix test cases, remove AppInsights ref from AutoML (#3329)

* [AutoML] disable netfx build leg for now (#3331)

* disable netfx build leg for now

* disable netfx build leg for now.

* [AutoML] Add AutoML XML documentation to all public members; migrate AutoML projects & tests into ML.NET solution; AutoML test fixes (#3351)

* [AutoML] Rev AutoML public API; add required native references to AutoML projects (#3364)

* [AutoML] Minor changes to generated project in CLI based on feedback (#3371)

* nitpicks for generated project

* revert back the target framework

* [AutoML] Migrate AutoML back to its own solution, w/ NuGet dependencies (#3373)

* Migrate AutoML back to its own solut…
@ghost ghost locked as resolved and limited conversation to collaborators Mar 31, 2022
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.
Labels
None yet
Projects
None yet
Development

Successfully merging this pull request may close these issues.

6 participants